-
Beihang University
- Shenzhen, China
-
02:35
(UTC +08:00) - https://zhoues.github.io/
Lists (19)
Sort Name ascending (A-Z)
💸 3D Assert
🌐 3D Vision
⏲️ 4D Vision
😃 Agent
⭐ awesome-paper-list
🦾 Bimanual Manipulation
👀 Ego
🤗 Embodied AI
🥇 Foundation Model
🎨 Image Generation
🤖 Minecraft Agent
Agent in Minecraft😆 MLLM & LLM
🚀 Navigation
🤯 Reasoning
🎇 Spatial Intelligence
🔧 Useful Tool
📹 Video
💯 VLM for Robotics
🎥 World Model
Stars
Skills for Real Engineers. Straight from my .agents directory.
Official Codebase for "Do as I Do: Dexterous Manipulation Data from Everyday Human Videos"
A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning
Provide with pre-build flash-attention 2 and 3 package wheels on Linux and Windows using GitHub Actions
A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
A paper list for Learning-based 3D Vision.
A practical toolkit for process-level robot evaluation with Process Reward Models (PRMs).
📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for streaming video.
Panoramic Affordance Prediction (PAP) (ECCV 2026)
[RSS 2026] Interactive World Simulator for Robot Policy Training and Evaluation
[CVPR 2025] VideoWorld is a simple generative model that learns purely from unlabeled videos—much like how babies learn by observing their environment.
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
[NeurIPS 2025 Spotlight] Towards Understanding Camera Motions in Any Video
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
[RSS 2026] Causal video-action world model for generalist robot control
Being-H is BeingBeyond's family of human-centric embodied foundation models.
SPAgent, a foundation agent for understanding, reasoning over, and operating within the physical and spatial world.
Orient Anything V2, NeurIPS 2025 Spotlight
[ICCV 2025] VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Official code for "TraceGen: World Modeling in 3D Trace-Space Enables Learning from Cross-Embodiment Videos" (CVPR 2026)
[NeurIPS 2025] Pixel-Perfect Depth
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation