Stars
[CVPR26] GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry
Official code for GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
CoTracker is a model for tracking any point (pixel) on a video.
Official implementation of Déjà View: Looping Transformers for Multi-View 3D Reconstruction
ViPE: Video Pose Engine for Geometric 3D Perception
Official implementation of paper "VLM³: Vision Language Models Are Native 3D Learners".
A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning
[CVPR'26] EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot Personalization
The LLVM for Robot Descriptions. A programmable IR engine to compose, validate, and compile URDF/XACRO/SRDF models from Python or Blender.
SOMA: From Surface Observations to Muscle Anatomy - 2026 - ECCV
Adds ability to download models from IKEA website
Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
[ICML 2026] 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
A tutorial and a set of tools to compute depth-from-stereo with Project Aria Gen2 devices. This includes stereo image rectification as well as disparity estimation
[CVPR 2026] UniCorrn: Unified Correspondence Transformer Across 2D and 3D
[NeurIPS 2025] Tracking and Understanding Object Transformations
[ICLR 2026] NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
Any4D: Unified Feed-Forward Metric 4D Reconstruction
Implementation of D4RT, Efficiently Reconstructing Dynamic Scenes, from Deepmind
pytorch implementation of "Efficiently Reconstructing Dynamic Scenes One 🎯 D4RT at a Time"
Awesome work on object 6 DoF pose estimation
[CVPR'26] TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens