-
Peking University
- BeiJing China
Stars
Intentional Updates for Streaming Reinforcement Learning
🎓 系统性大语言模型构建课程|🛠️ 覆盖预训练数据工程、Tokenizer、Transformer、MoE、GPU 编程 (CUDA/Triton)、分布式训练、Scaling Laws、推理优化及对齐 (SFT/RLHF/GRPO)|🚀 6 个渐进式作业 + 代码驱动,建立 LLM 全栈认知体系
A self-hosted ML coding practice platform. 68 problems from ReLU to flow matching — attention, training, RLHF, diffusion, and more. Instant feedback in the browser.
Vero: An Open RL Recipe for General Visual Reasoning
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
[ACL 2026] "OpenPhone: Mobile Agentic Foundation Models for AI Phone"
DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
[ICML 2026 Spotlight] On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
[CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
DigiData: Training and evaluating general-purpose mobile control agents
Tiny-FSDP, a minimalistic re-implementation of the PyTorch FSDP
Train a 1B LLM with 1T tokens from scratch by personal
[COLM 2025] Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
Tiny-DeepSpeed, a minimalistic re-implementation of the DeepSpeed library
A lightweight reinforcement learning framework that integrates seamlessly into your codebase, empowering developers to focus on algorithms with minimal intrusion.
[AAAI 2026] GUI-G²: Gaussian Reward Modeling for GUI Grounding
Scaling Preference Data Curation via Human-AI Synergy
slime is an LLM post-training framework for RL Scaling.
A simple framework to pre-train and fine-tune T5 model with pytorch-lightning and transformers
The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production.