Stars
iMac: Translating Actions into Motion and Contact Images for Embodied World Models
[Arxiv 2026] Toward the Whole Picture: Accumulative Fingerprint Mapping and Reconstruction for Small-Area Mobile Sensors
The official application of Identity-Consistent Multi-Pose Generation of Contactless Fingerprints
GesVLA: Gesture-Aware Vision-Language-Action Model with Embedded Representations
[CVPR 2026] AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
[ACL 2026 Poster] Code and Benchmark for "Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision"
RoboChallenge Inference example code
Official Hardware Codebase for the Paper "BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities"
[ICCV 2025] D^3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
[ICLR 2026] Code of "MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation"
Codes, datasets, and synthetic dataset generator about the paper "LiCamPose: Combining Multi-View LiDAR and RGB Cameras for Robust Single-Snapshot 3D Human Pose Estimation"
Nav-R1: Reasoning and Navigation in Embodied Scenes
[CoRL 2025] GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
[ICRA 2026] Official implementation of the paper: "StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling"
[ICCV 2025] IGL-Nav: Incremental 3D Gaussian Localization for Image-goal Navigation
RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. πππ
VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
[CoRL 2025] Repository relating to "TrackVLA: Embodied Visual Tracking in the Wild"
Official implementation of "OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning"
[CoRL25] GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
[ICML 2025 Oral] Official repo of EmbodiedBench, a comprehensive benchmark designed to evaluate MLLMs as embodied agents.
[WACV'24] TD3D: Top-Down Beats Bottom-Up in 3D Instance Segmentation
[RSS'25] This repository is the implementation of "NaVILA: Legged Robot Vision-Language-Action Model for Navigation"
This is a PyTorch implementation of MCLN proposed by our paper "Multi-branch Collaborative Learning Network for 3D Visual Grounding"(ECCV2024)