-
The University of Hong Kong
- Hong Kong SAR
-
18:45
(UTC +08:00) - jinghuahou7@gmail.com
Highlights
- Pro
Stars
RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
Welcome to GR00T Whole-Body Control (WBC)! This is a unified platform for developing and deploying advanced humanoid controllers. This includes: Decoupled WBC models used in NVIDIA Isaac-Gr00t, Gr0…
Vision as Unified Multimodal Generation
EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
[CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Models
An imitation learning stack for AgileX Piper arms and Cobot Magic systems—covering the full pipeline from hardware bring-up and teleop data collection to replay, LeRobot conversion, and policy infe…
ViPE: Video Pose Engine for Geometric 3D Perception
[CVPR 2026 (Highlight)] Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
[ECCV 2026] Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild
Cambrian-S: Towards Spatial Supersensing in Video
[ICLR 2026] FastVGGT: Fast Visual Geometry Transformer
[RSS 2025] "ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills"
[CVPR 2026] "E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training" official implementation.
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
[ICLR 2026] Trace Anything: Representing Any Video in 4D via Trajectory Fields
[CVPR'25 Oral] MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
[NeurIPS 2025] LabelAny3D: Label Any Object 3D in the Wild
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
[CVPR 2026] "GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation"
[CVPR 2026] DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning