Lists (8)
Sort Name ascending (A-Z)
Starred repositories
Code for "FACT: Failure-Aware Causal Training for World‑Action Models"
DeepSeek Harness: Everything is a Plugin.
SCoPE: Sightline-Coordinate Positional Encoding for 3D-Aware Video Generation
LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
Official Code of SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
Repository associated with paper titled "FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation", presented at ICML 2026.
Training and evaluation code for paper "Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs"
[CVPR 2026] Official implementation of "Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation"
This repo is the official implementation of "τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation".
[CVPR2026] Chain of World: World Model Thinking in Latent Motion
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
(ICML2026) Official implementation of VLANeXt.
Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation
INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models.
Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.
Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models
[ICML 2026] Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models
[CVPR 2026] Official implementation of "ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models"
Official project page and code repository for Enfold.
[Tech Report] Context Scaling: Scaling Properties of Text Conditioning in Visual Generation
Official repository of WCM, A World Critic Model for Vision-Language-Action Reinforcement Learning.
Implementation of RoboTTT proposed by Yunfan Jiang et al. of Stanford and Nvidia