Starred repositories
ScaRF-SLAM: Scale-Consistent Reconstruction with Feed-Forward Models and Classical Visual SLAM
Official implementation of CVPR2026 paper "AirSim360: A Panoramic Simulation Platform within Drone View"
Metaskill: A Meta-Skill for Autonomous AI Agent Team Generation
[ICRA'2024] Rethinking Imitation-based Planner for Autonomous Driving
Official support TITA Reinforcement Learning repo
A unified, agentic system for general-purpose robots, enabling multi-modal perception, mapping and localization, and autonomous mobility and manipulation, with intelligent interaction with users.
MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
NVIDIA Alpamayo 1 Nano is an open 10B reasoning VLA model for autonomous vehicles that pairs driving trajectories with Chain-of-Causation reasoning.
An open-source navigation stack based on Odin1.
Depth estimation by stereo fisheye camera (Cali Cam)
HoloMotion: A Foundation Model for Whole-Body Humanoid Control
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
Running VLA at 30Hz frame rate and 480Hz trajectory frequency
An invariant filter to fuse monocular/stereo visual-inertial-raw GNSS.
MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models
📖[IEEE Sensors Journal (JSEN) ] SuperVINS: A Real-Time Visual-Inertial SLAM Framework for Challenging Imaging Conditions (integrated deep learning features)
MATRiX is an advanced simulation platform that integrates MuJoCo, Unreal Engine 5, and CARLA to provide high-fidelity, interactive environments for robotics research.
[ICCV'25] Unified Open-World Segmentation with Multi-Modal Prompts
StreamingVLM: Real-Time Understanding for Infinite Video Streams
[ICLR 2026] Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
3DGS-to-PC: 3D Gaussian Splatting to Dense Point Clouds [3D-VAST: ICCVW 2025]
InteriorGS: 3D Gaussian Splatting Dataset of Semantically Labeled Indoor Scenes
[NeurIPS'25] Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Fully Open Framework for Democratized Multimodal Training