-
Google
- Switzerland
Stars
[ICLR2025 Spotlight๐ฅ] Official Implementation of TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
[ECCV 2024] Official Release of SILC: Improving vision language pretraining with self-distillation
[CVPRW-25 MMFM] Official repository of paper titled "How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs".
[ECCV2024 Oral๐ฅ] Official Implementation of "GiT: Towards Generalist Vision Transformer through Universal Language Interface"
[AAAI'25, CVPRW 2024] Official repository of paper titled "Learning to Prompt with Text Only Supervision for Vision-Language Models".
[ECCV'24] Official Implementation of SemiVL: Semi-Supervised Semantic Segmentation with Vision-Language Guidance
Language-Driven Semantic Segmentation
Convert Machine Learning Code Between Frameworks
Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
A concise but complete implementation of CLIP with various experimental improvements from recent papers
Code and data for the paper: "Paragraph-level Commonsense Transformers with Recurrent Memory"
The Curious Layperson: Fine-Grained Image Recognition without Expert Labels (BMVC 2021 best student paper)
A unified 3D Transformer Pipeline for visual synthesis
PyTorch CZSL framework containing GQA, the open-world setting, and the CGE and CompCos methods.
๐ A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
๐ Geometric Computer Vision Library for Spatial AI
Code base for the precision, recall, density, and coverage metrics for generative models. ICML 2020.
Pretrained bag-of-local-features neural networks