Stars
End-to-End Visual Language Navigation with Limited Sensing: A Survey
Official codebase for LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation
NaVIDA: Vision-Language Navigation with Inverse Dynamics Augmentation
Official Implementation for paper "Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm"
[CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.
[EMNLP 2025 Oral] Official codebase for Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors.
Official implementation of the paper: "ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation"
Structured Video Comprehension of Real-World Shorts
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. (CVPR 2025))
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
Visualizing the attention of vision-language models
My learning notes for ML SYS.
A curated list of awesome Vision-and-Language Navigation(VLN) resources (continually updated)
✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Machine Learning Engineering Open Book
From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
Lexical is an extensible text editor framework that provides excellent reliability, accessibility and performance.
[ACL 2024] A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
A trend starts from "Chain of Thought Prompting Elicits Reasoning in Large Language Models".
🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), ga…
✨✨Latest Advances on Multimodal Large Language Models
Mini Database System using B+ Tree in C++ (Simple & Self-Explanatory Code)
A seamless, speedy, syncing markdown editor