Skip to content
View ydsf16's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report ydsf16

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[ICCV 2023] ProPainter: Improving Propagation and Transformer for Video Inpainting

Python 6,894 815 Updated Feb 19, 2025

EgoVerse: Egocentric Data for Robot Learning from Around the World

Jupyter Notebook 524 46 Updated Aug 18, 2026

[ECCV 2026] SLAM-Former: Putting SLAM into One Transformer

Python 500 11 Updated Aug 10, 2026

Deep Patch Visual Odometry/SLAM

C++ 1,103 165 Updated Oct 12, 2024

[CVPR 2026 Oral] VGGT Omega

Python 4,060 282 Updated Aug 18, 2026
Python 2,160 409 Updated Jul 23, 2024

🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning

Python 26,754 5,417 Updated Aug 19, 2026

Advancing Open-source World Models

Python 4,374 396 Updated Jul 9, 2026

A real-time system that simultaneously captures human pose, reconstructs the scene in sparse 3D points, and localizes the human in the scene with 6 IMUs and a body-worn phone camera

C++ 114 21 Updated May 8, 2025

A real-time motion capture system that estimates poses and global translations using only 6 inertial measurement units

Python 449 84 Updated May 8, 2025

Official Repo for Qwen-RobotNav

150 9 Updated Jun 30, 2026

A Modular and Multi-Modal Mapping Framework

C++ 2,866 749 Updated May 31, 2024

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,804 1,831 Updated Jan 30, 2026

从0开始学习多模态大模型,从概念、架构、训练到部署,系统搭建你的第一套 MLLM 知识体系

142 9 Updated Jul 9, 2026

The best ChatGPT that $100 can buy.

Python 57,307 7,957 Updated Aug 2, 2026

Coding a Multimodal (Vision) Language Model from scratch in PyTorch with full explanation: https://www.youtube.com/watch?v=vAmKB7iPkWw

Python 630 105 Updated Dec 6, 2024

Self-supervised learning for spatial perception

Python 916 43 Updated Jul 8, 2026

Video+code lecture on building nanoGPT from scratch

Python 5,431 879 Updated Aug 13, 2024
Python 1 Updated Jul 6, 2026
Swift 23 Updated Jul 22, 2026

Real-time AI assistant for Meta Ray-Ban smart glasses -- voice + vision + agentic actions via Gemini Live and OpenClaw

TypeScript 2,517 492 Updated Aug 19, 2026

[CVPR 2025] EgoLife: Towards Egocentric Life Assistant

Python 457 20 Updated Mar 19, 2025
Python 1,442 118 Updated Feb 12, 2026

Source code for the ECCV 2022 paper "Benchmarking Localization and Mapping for Augmented Reality".

Python 440 48 Updated Oct 27, 2025

AirSteady codes

C++ 6 Updated Apr 19, 2026

个人主站

1 Updated Apr 18, 2026

空间智能合集

Python 9 1 Updated Apr 14, 2026

GIM: Learning Generalizable Image Matcher From Internet Videos (ICLR 2024 Spotlight)

Python 902 61 Updated Aug 3, 2025
Next