-
Computer Vision Lab, POSTECH
- kdwonn.github.io
Stars
[Tech Report] Democratizing the Training of Video World Models from Scratch. 🔥 🔥 🔥
FlowWM stochastic world modeling via flow matching in DINOv3 feature space, with the FuturePerception (Waymo) benchmark.
A Curated List of Awesome Video World Models with AR Diffusion: Covering Algorithms, Applications, and Infrastructure, Aimed at Serving as a Comprehensive Resource for Researchers, Practitioners, a…
Open-source Dreamer world-model implementation in JAX
Unofficial implementation of the Dreamer 4 world model in PyTorch.
Code for MIRA: Multiplayer Interactive World Models with Representation Autoencoders
Dimensional is the agentic operating system for physical space. Command humanoids, quadrupeds, drones, and other hardware platforms in natural language and build multi-agent systems that work seaml…
DROID Policy Learning and Evaluation
A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Official implementation of "Geometric Action Model for Robot Policy Learning"
Official implementation of "Turning Video Models into Generalist Robot Policies"
ABC: Scalable Behavior Cloning with Open Data, Training, and Evaluation
Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
Use Lerobot to collect piper robot arm data, and perform training and reasoning 使用lerobot采集piper机械臂数据,并训练和推理
[CVPR 2026] Affostruction: 3D Affordance Grounding with Generative Reconstruction
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
A Minimalist, Batteries-included Repository for Advancing World Model Science.
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR 2025.
A curated list of papers and selected technical blogs on Loop Models.
Unfied World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
[RSS 2026] LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
An offline-first scientific writing workspace powered by Claude. LaTeX + Python + 100+ scientific skills all running locally.
[CVPR 2024] Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform