Stars
[CoRL2026] OpenHLM: An Empirical Recipe for Whole-Body Humanoid Loco-Manipulation
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control [Under Review]
Native-video memory for vision-language-action models, using timestamped visual history and exact streaming inference for long-horizon robot manipulation.
Official repository for "OpenWAM: An Open, Modular Exploration Towards Systematic World–Action Model Pretraining".
Extensible video-action world models for robot learning
[ICRA 2026] GMR: General Motion Retargeting. Retarget human motions into diverse humanoid robots in real time on CPU. Retargeter for TWIST.
[CoRL 2026] EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment
[RSS'26] Welcome to Psi-Zero, a Humanoid VLA towards Universal Humanoid Intelligence.
[RSS 2026] The first framework enabling humanoid robots to learn whole-body loco-manipulation from egocentric human demos
[ICLR 2026] Towards Unified Latent VLA for Whole-body Loco-manipulation Control
SCAN-Planner: Spatial Collision-Aware Local planning for Route-Guided Long-Range Quadruped Navigation
Paper list in the survey paper: Challenges and Opportunities of Using Deep Learning in Safe-Critical Robotic Manipulator Planning
Flex-π: A multi-stream world-action model with compute flexibility: one checkpoint that deploys as a VLA, a full world model, or anything in between.
Trajectory Optimization Motion Planner for ROS
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
The absolute trainer to light up AI agents.
M2T2: Multi-Task Masked Transformer for Object-centric Pick and Plac
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
Addressing the Orchestration Gap in Generalist Robots via Physical Agency