Starred repositories
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping [CVPR 2025]
[CVPR 2021 Oral] Reconstructing 3D Human Pose by Watching Humans in the Mirror
Code release for paper "Reconstructing People, Places, and Cameras", In CVPR 2025 (Highlight)
Shareable model files (urdf/sdf + meshes, etc) for our robotics projects
😎 A curated list of awesome collision detection libraries and resources
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
[Humanoids 2025 Oral] Universal Humanoid Robot Pose Learning from Internet Human Videos
[RSS 2025] AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control
[ICCV 2025] ZeroStereo: Zero-Shot Stereo Matching from Single Images
Single-file implementation to advance vision-language-action (VLA) models with reinforcement learning.
ROEVO: Robust Organized Edge Feature-based Visual Odometry Using RGB-D Cameras
Official Code of Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
RynnEC: Bringing MLLMs into Embodied World
[CVPR25 Oral (Top 3.3%)] Official code for paper "Reconstructing Humans with a Biomechanically Accurate Skeleton".
[TPAMI 2025 & CVPR 2023] IGEV++: Iterative Multi-range Geometry Encoding Volumes for Stereo Matching
Implementation of Dex1B: Learning with 1B Demonstrations for Dexterous Manipulation, from Ye et al of UCSD
Official implementation of ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.
Controlling diverse robots by inferring jacobian fields with deep networks! Let's make robots understand their bodies!
[RSS 2025] ROMAN: a view-invariant global localization method that matches objects from different robot views for reliable pose estimation even when a scene is observed from opposite views
【CVPR 2025 Highlight】MonSter: Marry Monodepth to Stereo Unleashes Power
NVIDIA NeMo Canary-Qwen-2.5B is an English speech recognition model that achieves state-of-the art performance on multiple English speech benchmarks.
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.