-
Samsung Research
- Cambridge, UK
-
09:58
(UTC +01:00) - https://abaldrati.github.io
- in/alberto-baldrati
- @A_Baldrati
- https://scholar.google.com/citations?user=I1jaZecAAAAJ&hl=en
Stars
A paper list of some recent works about Token Compress for Vit and VLM
[CVPR 2026 Highlight] A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
[CVPR 2026 Highlight] ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
Official repo for "Let ViT Speak: Generative Language-Image Pre-training"
[CVPR 2026] Elucidating the SNR-t Bias of Diffusion Probabilistic Models
[CVPR 2026] - IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders [Technical Report]
[CVPR26 Findings] PEPR: Privileged Event-based Predictive Regularization for Domain Generalization
This is a research project to efficiently compress diffusion models, accepted at AAAI 2026.
[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
[WACV 2026] Official implementation of the paper: “CountingDINO: A Training-free Pipeline for Exemplar-based Class-Agnostic Counting”
[CVPR'26 Demo] Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
[ICLR 2026] - Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
Code for NeurIPS 2025 paper - Covariances for Free: Exploiting Mean Distributions for Training-free Federated Learning
[ICML 2025] No Task Left Behind: Isotropic Model Merging with Common and Task-Specific Subspaces (official repository)
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
This repository contains code for the paper "Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training" by T. Bonnaire, R. Urfin, G. Biroli and M. Mézard.
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
Official implementation of the paper: "FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models"
Official inference repo for FLUX.2 models
[AAAI2026] Mitigating Negative Flips via Margin Preserving Training
[NeurIPS 2025] Official implementation of "Instance-Level Composed Image Retrieval".
[ICCV 2025] What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
[TMLR] Public code repo for paper "A Single Transformer for Scalable Vision-Language Modeling"