-
Brown University
- Providence, RI
Highlights
- Pro
Starred repositories
ABC: Scalable Behavior Cloning with Open Data, Training, and Evaluation
Unfied World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
Being-H is BeingBeyond's family of human-centric embodied foundation models.
A Minimalist, Batteries-included Repository for Advancing World Model Science.
An open-source implementation of Multitask Diffusion Transformer (DiT) Policy for robot manipulation as seen on Boston Dynamics Atlas Humanoid and in "A Careful Examination of Large Behavior Models…
One framework to evaluate any VLA model on any robot simulation benchmark.
[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
[NeurIPS 2025 Oral] Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
[ICML 2026] Orienting Latent Actions for Video World Modeling
[ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
[RSS 2025] Learning to Act Anywhere with Task-centric Latent Actions
[ICLR 2025] LAPA: Latent Action Pretraining from Videos
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
A optimized PyTorch framework for behavior cloning with flow related generative models.
Code for "Evaluating Robot Policies in a World Model".
XLeRobot: Practical Dual-Arm Mobile Home Robot for $660
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
[ICLR 2026] Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
A procedural Blender pipeline for photorealistic training image generation
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
A general framework for inference-time scaling and steering of diffusion models with arbitrary rewards.
Paper list for Efficient Reasoning.
[NeurIPS 2025 Spotlight] LLM post-training suite — featuring ReasonFlux, ReasonFlux-PRM, and ReasonFlux-Coder.
Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
code for "CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models"