-
BUAA
- Beijing, China
- https://www.zhihu.com/people/DongShengYang/columns
Lists (3)
Sort Name ascending (A-Z)
Stars
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
A real-time system that simultaneously captures human pose, reconstructs the scene in sparse 3D points, and localizes the human in the scene with 6 IMUs and a body-worn phone camera
A real-time motion capture system that estimates poses and global translations using only 6 inertial measurement units
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Coding a Multimodal (Vision) Language Model from scratch in PyTorch with full explanation: https://www.youtube.com/watch?v=vAmKB7iPkWw
Self-supervised learning for spatial perception
Video+code lecture on building nanoGPT from scratch
Real-time AI assistant for Meta Ray-Ban smart glasses -- voice + vision + agentic actions via Gemini Live and OpenClaw
[CVPR 2025] EgoLife: Towards Egocentric Life Assistant
Source code for the ECCV 2022 paper "Benchmarking Localization and Mapping for Augmented Reality".
GIM: Learning Generalizable Image Matcher From Internet Videos (ICLR 2024 Spotlight)
open Multi-View Stereo reconstruction library