Stars
RPent: Agentic Infrastructure for the Physical World
[RSS 2026] Causal video-action world model for generalist robot control
The official Implementation of GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation
Elemental Diagnosis of Generalist Mobile Manipulation Policies
Machine Learning Engineering Open Book
The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that sho…
[ECCV 2026] Official Code of "Distribution Matching Distillation Meets Reinforcement Learning"
Official PyTorch implementation of FB-BEV & FB-OCC - Forward-backward view transformation for vision-centric autonomous driving perception
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
Support robosense LiDAR including M1, E1R, and Airy
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
A high-throughput and memory-efficient inference and serving engine for LLMs
[ICLR 2026] From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
A MemAgent framework that can be extrapolated to 3.5M, along with a training framework for RL training of any agent workflow.
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
[RSS'25] This repository is the implementation of "NaVILA: Legged Robot Vision-Language-Action Model for Navigation"
[CVPR 2025 Highlight] Towards Autonomous Micromobility through Scalable Urban Simulation
Enhancing Autonomous Driving Systems with On-Board Deployed Large Language Models
Train transformer language models with reinforcement learning.