Skip to content
View ABaldrati's full-sized avatar

Organizations

@miccunifi

Block or report ABaldrati

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A paper list of some recent works about Token Compress for Vit and VLM

944 44 Updated Jul 27, 2026
55 Updated Jul 26, 2026

[CVPR 2026 Highlight] A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

Python 219 4 Updated Jul 17, 2026

[CVPR 2026 Highlight] ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

Python 28 Updated Jul 13, 2026

Official repo for "Let ViT Speak: Generative Language-Image Pre-training"

Python 133 4 Updated Jun 10, 2026

timm, evolved

Python 60 Updated May 28, 2026
Python 18 Updated Apr 17, 2026

[CVPR 2026] Elucidating the SNR-t Bias of Diffusion Probabilistic Models

Python 120 3 Updated Apr 20, 2026

[CVPR 2026] - IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment

Python 29 3 Updated May 29, 2026

🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.

Jupyter Notebook 4,427 388 Updated Jul 31, 2026

Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders [Technical Report]

Jupyter Notebook 206 13 Updated Mar 30, 2026

[CVPR26 Findings] PEPR: Privileged Event-based Predictive Regularization for Domain Generalization

HTML 6 Updated May 23, 2026

This is a research project to efficiently compress diffusion models, accepted at AAAI 2026.

Python 6 Updated Mar 5, 2026

[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

Python 695 21 Updated May 23, 2026

[WACV 2026] Official implementation of the paper: “CountingDINO: A Training-free Pipeline for Exemplar-based Class-Agnostic Counting”

Jupyter Notebook 64 5 Updated Jun 22, 2026

[CVPR'26 Demo] Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device

Python 154 16 Updated Apr 13, 2026

[ICLR 2026] - Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery

Python 23 Updated Mar 18, 2026

Code for NeurIPS 2025 paper - Covariances for Free: Exploiting Mean Distributions for Training-free Federated Learning

Python 9 Updated Apr 16, 2026

Revisiting Multi-Task Visual Representation Learning

22 Updated Jan 21, 2026

[ICML 2025] No Task Left Behind: Isotropic Model Merging with Common and Task-Specific Subspaces (official repository)

Python 46 3 Updated Aug 7, 2025

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Python 4,578 349 Updated Jan 14, 2026

This repository contains code for the paper "Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training" by T. Bonnaire, R. Urfin, G. Biroli and M. Mézard.

Python 83 9 Updated Nov 27, 2025

Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models

Python 9 Updated Mar 14, 2026

Official implementation of the paper: "FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models"

Python 1,012 51 Updated May 27, 2026

Official inference repo for FLUX.2 models

Python 2,592 184 Updated Mar 12, 2026

[AAAI2026] Mitigating Negative Flips via Margin Preserving Training

Python 5 Updated Nov 16, 2025

[NeurIPS 2025] Official implementation of "Instance-Level Composed Image Retrieval".

Python 53 1 Updated Dec 22, 2025

[ICCV 2025] What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

Python 16 Updated Nov 3, 2025

[TMLR] Public code repo for paper "A Single Transformer for Scalable Vision-Language Modeling"

Jupyter Notebook 150 4 Updated Nov 14, 2024
Next