Stars
Post-training with Tinker
[NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL
[2025] Efficient Vision Language Models: A Survey
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
Towards Efficient Multimodal Large Language Models: A Survey on Token Compression
[Up-to-date] Large Language Model Agent: A Survey on Methodology, Applications and Challenges
A course in reinforcement learning in the wild
👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
collection of diffusion model papers categorized by their subareas
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation (ICCV'25)
A compilation of the best multi-agent papers
Evolve your language agent with Agentic Context Engineering (ACE)
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
[ICCV'23 Main Track, WECIA'23 Oral] Official repository of paper titled "Self-regulating Prompts: Foundational Model Adaptation without Forgetting".
[ICLR 2026] The offical Implementation of "Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model"
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
EmoBench-M: A benchmark for evaluating Emotional Intelligence in Multimodal Large Language Models (MM 2026)
🔥 🔥 🔥 A paper list of some recent Computer Vision(CV) works
A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.
A toolbox for skeleton-based action recognition.
A flexible and extensible framework for gait recognition. You can focus on designing your own models and comparing with state-of-the-arts easily with the help of OpenGait.
A curated list of Gait Recognition and related area resource
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A PyTorch implementation of the Transformer model in "Attention is All You Need".
Homepage for STAT 157 at UC Berkeley
deep learning for image processing including classification and object-detection etc.