Stars
Our inference and training framework to run on the Cosmos Models
[CVPR'25 Oral] Official implementation for "DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models"
[SIGGRAPH Asia 2026] Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR 2025.
NVIDIA Alpamayo 2 Super is an open 34B multi-task foundation model designed to supercharge autonomous vehicle development.
The first multiplayer video world model in Minecraft
Code for "Scaling Language-Free Visual Representation Learning" paper (Web-SSL).
Our insights of Openpilot, a deepdive project on it
InstantNuRec: Feed-Forward 3D Gaussian Reconstruction from Driving Logs
A curated list of materials on AI efficiency
N₀-TWAM: A Tactile-Native World Action Model for Contact-Rich Manipulation
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
[SIGGRAPH Asia 2026] UniMate: One Unified Model to Animate Diverse Skeletons
[CVPR 2026] PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
[CVPR 2026] Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction
[CVPR 2026] The official PyTorch implementation of the "Vision Transformer Needs More Than Registers".
A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
Official implementation of ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation (SIGGRAPH 2026).