LLM 学习笔记:Transformer 架构、强化学习 (RLHF/DPO/PPO)、分布式训练、推理优化。含完整数学推导与Slides。
-
Updated
Feb 28, 2026 - TeX
LLM 学习笔记:Transformer 架构、强化学习 (RLHF/DPO/PPO)、分布式训练、推理优化。含完整数学推导与Slides。
Bilingual graduate textbook on LLM post-training: SFT, RLHF, DPO, GRPO, RLVR, reasoning, agents, systems, evaluation, and safety.
Exact finite-group identity behind GRPO reward standardization, unifying GRPO / Dr. GRPO / DAPO for RLVR and LLM reasoning. Paper + code.
Comprehensive university-level study guide for LLMs, transformers, RLHF, and generative AI. 335+ pages (actively expanding), 47 visualizations, 4 notebooks. Regular updates with enhanced sections, new implementations, and expanded coverage.
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
AI Text Slop: A Quantitative Study of Stylistic Convergence Across Six Language Models in Japanese Technical Writing
To associate your repository with the rlhf topic, visit your repo's landing page and select "manage topics."