Stars
TTRV: Test-Time Reinforcement Learning for Vision–Language Models (CVPR 2026)
VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.
Repository for the paper: Teaching VLMs to Localize Specific Objects from In-context Examples
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…
ImageBind One Embedding Space to Bind Them All
Official repository for the MMFM challenge
Official inference library for Mistral models
Video Test-Time Adaptation for Action Recognition (CVPR 2023)
Code for the paper: Rotating Features for Object Discovery
Cleaner and Formatter for BibTeX files
Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"
Automated dense category annotation engine that serves as the initial semantic labeling for the Segment Anything dataset (SA-1B).
Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/sp…
The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
Official implementation of Cold-Diffusion for different transformations in pytorch.
Implementation of Denoising Diffusion Probabilistic Model in Pytorch
Unofficial implementation of Palette: Image-to-Image Diffusion Models by Pytorch
code for deep learning courses
Official PyTorch implementation of ODISE: Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models [CVPR 2023 Highlight]
The simplest, fastest repository for training/finetuning medium-sized GPTs.
EVA Series: Visual Representation Fantasies from BAAI
Test-time Prompt Tuning (TPT) for zero-shot generalization in vision-language models (NeurIPS 2022))
3D point cloud datasets in HDF5 format, containing uniformly sampled 2048 points per shape.
pytorch based implementation faster rcnn
A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training
Repo for "Benchmarking Robustness of 3D Point Cloud Recognition against Common Corruptions" https://arxiv.org/abs/2201.12296
Refine high-quality datasets and visual AI models