Highlights
- Pro
Stars
The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that sho…
Official codes for paper: 3DGS-DET: Empower 3D Gaussian Splatting with Boundary Guidance and Box-Focused Sampling for Indoor 3D Object Detection
😎 Awesome lists of papers and codes about Large Vision-Language Models
😎 Awesome lists of papers and codes about open-vocabulary perception, including both 3D and 2D
Awesome Data-Driven Autonomous Driving Solutions. Also the official repository of our survey paper: Data-Centric Evolution in Autonomous Driving: A Comprehensive Survey of Big Data System, Data Min…
Emu Series: Generative Multimodal Models from BAAI
[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
(TPAMI 2024) A Survey on Open Vocabulary Learning
Official code for NeurIPS2023 paper CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection and TPAMI 2025 paper CoDAv2
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
🧀 Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs".
[ECCV 2024] Official implementation of the paper "Semantic-SAM: Segment and Recognize Anything at Any Granularity"
(ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
This is the official code for MobileSAM project that makes SAM lightweight for mobile applications and beyond!
A collection of resources and papers on Diffusion Models
Segment Anything in High Quality [NeurIPS 2023]
A collection of papers on the topic of ``Computer Vision in the Wild (CVinW)''
A curated list of foundation models for vision and language tasks
[ICRA 2022] Towards Scale Consistent Monocular Visual Odometry by Learning from the Virtual World
[IJCV 2022] Information-Theoretic Odometry Learning
MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.
The official repo for [AAAI 2024] "SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection""
The official repo for [TPAMI'23] "Vision Transformer with Quadrangle Attention"
Official implementation of "Composer: Creative and Controllable Image Synthesis with Composable Conditions"
Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.