Stars
[ECCV 2026] Fast SAM 3D Body: Accelerating SAM 3D Body for Real-Time Full-Body Human Mesh Recovery
SuperMap is a living spatial memory for embodied AI — it perceives the world, remembers its evolution, and supports reasoning and action.
Official PyTorch implementation of "InterHand2.6M: A Dataset and Baseline for 3D Interacting Hand Pose Estimation from a Single RGB Image", ECCV 2020
Large dataset of hand-object contact, hand- and object-pose, and 2.9 M RGB-D grasp images.
HaMeR: Reconstructing Hands in 3D with Transformers
HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
VPoser: Variational Human Pose Prior
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Models
Vision as Unified Multimodal Generation
WiLoR: End-to-end 3D hand localization and reconstruction in-the-wild
[CVPR'26 Highlight] Official Repository of CVPR2026 Paper: ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions
Official repo for Nymeria and NymeriaPlus datasets.
RealSee3D: A multi-view RGB-D dataset combining real-world captures and procedurally generated scenes, with extensible annotations for diverse 3D vision research.
[ECCV 2026] Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
[ICCV2023] EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding
Official implementation of ECCV24 paper "SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding"
High-fidelity world models for general embodied intelligence, such as data engines and world simulators.
Code release for "Omni3D A Large Benchmark and Model for 3D Object Detection in the Wild"
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
[3DV 2025 Oral]: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
让每一次引用都成为可解释的影响力 Turning Every Citation into Explainable Impact
PhotoFlow: Agentic 3D Virtual Photography Missions