Highlights
- Pro
Stars
Reproducible navigation benchmarking & residual RL fine-tuning for Wanderland
Reimplementation of LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
Official JAX implementation of End-to-End Test-Time Training for Long Context
A ROS driver for Insta360 cameras, enabling real-time image capture, processing, and publishing in ROS environments.
Official PyTorch implementation of One-Minute Video Generation with Test-Time Training
[CVPR 2026] Thinking in 360°: Humanoid Visual Search in the Wild
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations https://video-prediction-policy.github.io
Official Implementation of Paper Transfer between Modalities with MetaQueries
Cambrian-S: Towards Spatial Supersensing in Video
XLeRobot: Practical Dual-Arm Mobile Home Robot for $660
Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Lear…
VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments
[CVPR2025] PyTorch-based reimplementation of CrossFlow, as proposed in 'Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution'
This repo contains the code for 1D tokenizer and generator
Official repo for From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Lets make video diffusion practical!
RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
This package contains the original 2012 AlexNet code.
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
CoTracker is a model for tracking any point (pixel) on a video.
[CVPR 2025] Magma: A Foundation Model for Multimodal AI Agents
Official implementation of Continuous 3D Perception Model with Persistent State