Starred repositories
[CVPR 2026 Oral] "MARCO: Navigating the Unseen Space of Semantic Correspondence"
A data collection and processing pipeline for animal video, annotations include mask, keypoint, depth, occlusion, etc. Suitable for 3D/4D reconstruction, tracking, pose prediction, etc.
[ICLR 2026 oral] Official code for VIST3A: Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
🏂 Training-Free Human Mesh Recovery from Videos, based on SAM-3, Diffusion-VAS, and SAM-3D-Body.
Native and Compact Structured Latents for 3D Generation
Official implementation of the 2024 ECCV paper SHIC: Shape-Image Correspondences with no Keypoint Annotation
[ECCV 2024] Official implementation of the paper "X-Pose: Detecting Any Keypoints"
[CVPR 2026 Findings] TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
Muti-human Interactive Talking Dataset
Official implementation of "MoMask: Generative Masked Modeling of 3D Human Motions (CVPR2024)"
The ultimate training toolkit for finetuning diffusion models
Controllable video and image Generation, SVD, Animate Anyone, ControlNet, ControlNeXt, LoRA
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
[CVPR 2024] "LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning"; an interactive Large Language 3D Assistant.
Official Implementation of Diffusion Step Annealing (DiSA) in Autoregressive Image Generation
This repository contains an implementation for performing 3D animal (quadruped) reconstruction from a monocular image or video. The system adapts the pose (limb positions) and shape (animal type/he…
[IJCV 2022] Bridging Composite and Real: Towards End-to-end Deep Image Matting
Official implementation of DeepLabCut: Markerless pose estimation of user-defined features with deep learning for all animals incl. humans
official implementation of paper "Tactile DreamFusion: Exploiting Tactile Sensing for 3D Generation"
2027 AI/ML internship & new graduate job list updated daily
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
[CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
[CVPR2024 (Highlight)] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D. Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC
[3DV 2025 Best Paper] We present Object Images (Omages): An homage to the classic Geometry Images.
[ECCV 2024] Code for VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
[ICLR 2025] HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models