Highlights
- Pro
Stars
[Tech Report] Context Scaling: Scaling Properties of Text Conditioning in Visual Generation
HY-SOAR:Self-Correction for Optimal Alignment and Refinement in Diffusion Models
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
Official Implemenation for RAEv2: Improved Baselines with Representation Autoencoders
Official repository of LIBERO-plus, a generalized benchmark for in-depth robustness analysis of vision-language-action models.
[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Official PyTorch Implementation of "Flow Map Distillation Without Data"
This is the official code repo for DiT4DiT, a Vision-Action-Model (VAM) framework that combines video generation model with flow-matching-based action prediction for generalizable robotic manipulat…
Video-Action Models for Generalizable Robot Control Beyond VLAs
VLS: Steering Pretrained Robot Policies via Vision–Language Models
[ECCV 2026] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
Single-stage End-to-End Training for Tokenization and Generation
🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
Official Code Repository for the paper "Score-based Generative Modeling of Graphs via the System of Stochastic Differential Equations" (ICML 2022)
GigaWorld-Policy: An Efficient Action-Centered World–Action Model
An end-to-end open ecosystem for robot learning
Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
[RSS 2026] Causal video-action world model for generalist robot control
ACTSmooth extends ACT with prefix conditioning and async inference, eliminating inter-chunk discontinuities and inference latency stalls.
Unfied World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
[ICLR 2026] Official implementation for What matters for Representation Alignment: Global Information or Spatial Structure?
The first multiplayer video world model in Minecraft
(ICML2026) Official implementation of VLANeXt.