Stars
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
Official Implementation of SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models (ICML'26 Spotlight)
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
A general framework for inference-time scaling and steering of diffusion models with arbitrary rewards.
Official repo for Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models (ICML 2026 Spotlight)
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
[ECCV 2026] WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
Code for <Environmental Change Detection for Real-World Change Analysis> in ECCV 2026
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
Code for 'Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold' (ECCV 2026)
[ECCV 2026] SAM2Matting: Generalized Image and Video Matting
From a single casual image to a visually consistent and physically stable interactive 3D scene.
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
[CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
[CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image (CVPR 2026)
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
Official implementation of paper "VLM³: Vision Language Models Are Native 3D Learners".
A curated list of awesome 3D object and scene generation papers.
[CVPR2026] Detect Anything via Next Point Prediction