Stars
ViPE: Video Pose Engine for Geometric 3D Perception
[🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s Multimodal Intelligence team.
Reference PyTorch implementation and models for DINOv3
[ICRA 2026] RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
Hierarchical Reasoning Model Official Release
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
[IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
Ever used asyncio and wished you hadn't? A tiny (~400 lines) event loop for Python.
Label Studio is a multi-type data labeling and annotation tool with standardized output format
[ICRA 2025 & IJRR 2026] Gaussian-LIC2: LiDAR-Inertial-Camera Gaussian Splatting SLAM (in Real Time)
Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
color-icon-matrix barcodes. Proof of concept implementation.
PyTorch Lightning + Hydra. A very user-friendly template for ML experimentation. ⚡🔥⚡
Official inference repo for FLUX.1 models
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
[CVPR 2025 Oral & Best Paper Finalist] Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
PyTorch implementation of MeanFlow & iMF (one-step generative modeling).
High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
Solve Visual Understanding with Reinforced VLMs
[ICCV 2025] Official code of "ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation"
[ECCV 2024] Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning