Stars
Official implementation of "SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation".
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
[NeurIPS 2025] A Practical Guide for Incorporating Symmetry in Diffusion Policy
egocentric humanoid manipulation benchmark
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
[CVPR 2025 Highlight] Truncated Diffusion Model for Real-Time End-to-End Autonomous Driving
Various retargeting optimizers to translate human hand motion to robot hand motion.
VLA-0: Building State-of-the-Art VLAs with Zero Modification
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
[IEEE T-PAMI 2024] All you need for End-to-end Autonomous Driving
The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that sho…
Dexbotic: Open-Source Vision-Language-Action Toolbox
RoboChallenge Inference example code
Reference PyTorch implementation and models for DINOv3
[ICML 2025] Official implementation of Spherical Diffusion Policy: A SE(3) Equivariant Visuomotor Policy with Spherical Fourier Representation
[3DV 2025] Efficient Continuous Group Convolutions for Local SE(3) Equivariance in 3D Point Clouds
Testing flow matching in Euclidean space and Lie groups.
PyTorch implementation of MeanFlow & iMF (one-step generative modeling).
[NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL
A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities.
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
[ICML 2026] RoboTwin 2.0 Offical Code Repo