Skip to content
View ydsf16's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report ydsf16

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[CVPR 2026 Oral] VGGT Omega

Python 3,902 261 Updated Jul 15, 2026
Python 2,135 406 Updated Jul 23, 2024

🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning

Python 26,507 5,341 Updated Aug 8, 2026

Advancing Open-source World Models

Python 4,346 396 Updated Jul 9, 2026

A real-time system that simultaneously captures human pose, reconstructs the scene in sparse 3D points, and localizes the human in the scene with 6 IMUs and a body-worn phone camera

C++ 114 21 Updated May 8, 2025

A real-time motion capture system that estimates poses and global translations using only 6 inertial measurement units

Python 450 84 Updated May 8, 2025

Official Repo for Qwen-RobotNav

140 9 Updated Jun 30, 2026

A Modular and Multi-Modal Mapping Framework

C++ 2,863 748 Updated May 31, 2024

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,745 1,831 Updated Jan 30, 2026

从0开始学习多模态大模型,从概念、架构、训练到部署,系统搭建你的第一套 MLLM 知识体系

122 8 Updated Jul 9, 2026

The best ChatGPT that $100 can buy.

Python 57,052 7,907 Updated Aug 2, 2026

Coding a Multimodal (Vision) Language Model from scratch in PyTorch with full explanation: https://www.youtube.com/watch?v=vAmKB7iPkWw

Python 628 105 Updated Dec 6, 2024

Self-supervised learning for spatial perception

Python 898 39 Updated Jul 8, 2026

Video+code lecture on building nanoGPT from scratch

Python 5,412 878 Updated Aug 13, 2024
Python 1 Updated Jul 6, 2026
Swift 23 Updated Jul 22, 2026

Real-time AI assistant for Meta Ray-Ban smart glasses -- voice + vision + agentic actions via Gemini Live and OpenClaw

TypeScript 2,491 486 Updated Aug 6, 2026

[CVPR 2025] EgoLife: Towards Egocentric Life Assistant

Python 451 19 Updated Mar 19, 2025
Python 1,438 117 Updated Feb 12, 2026

Source code for the ECCV 2022 paper "Benchmarking Localization and Mapping for Augmented Reality".

Python 440 48 Updated Oct 27, 2025

AirSteady codes

C++ 6 Updated Apr 19, 2026

个人主站

1 Updated Apr 18, 2026

空间智能合集

Python 9 1 Updated Apr 14, 2026

GIM: Learning Generalizable Image Matcher From Internet Videos (ICLR 2024 Spotlight)

Python 902 61 Updated Aug 3, 2025

open Multi-View Stereo reconstruction library

C++ 4,080 983 Updated Aug 3, 2026

An open library of computer vision algorithms

C 1,648 619 Updated Aug 25, 2022

[TRO 2025] AirVO upgrades to AirSLAM

C++ 1,188 176 Updated Nov 19, 2025
Next