Skip to content
View KMnO4-zx's full-sized avatar
🎯
Time is all you need.
🎯
Time is all you need.
  • Beijing, China
  • 12:18 (UTC +08:00)

Highlights

  • Pro

Block or report KMnO4-zx

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
KMnO4-zx/README.md
KMnO4-zx banner


靡不有初,鲜克有终。

Zhihu Xiaohongshu Email Profile visitors

About Me

  • 🔬 Open-source developer and researcher focused on Agent RL, LLM post-training, and AI agents.
  • 🌟 Developer and maintainer of open-source projects with 100,000+ GitHub stars in total.
  • 🧪 Researcher at Emotion Machine Lab, also working on growth and developer/community operations.
  • 📫 Reach me at kmno4-song@foxmail.com.
GitHub contribution grid snake animation

Research Interests

  • Agent RL — reinforcement learning for reasoning, tool use, and multi-turn agents.
  • LLM Post-training — alignment, feedback mechanisms, and efficient post-training methods.
  • AI Agents — agentic systems, research agents, and evaluation.

Open Source Experience

Maintainer & Creator

  • Happy-LLM — A from-scratch guide to LLM fundamentals and implementation. Happy-LLM stars
  • self-llm — A practical guide to deploying and fine-tuning open-source LLMs. self-llm stars
  • tiny-universe — Hands-on implementations of RAG, agents, evaluation, and other LLM systems. tiny-universe stars
  • llm-agent-rl-lab — Reproductions and studies of RL algorithms for LLM agents. llm-agent-rl-lab stars
  • paper-insight — An AI-powered platform for paper analysis and research workflows. paper-insight stars
  • huanhuan-chat — A ChatGLM-based character chatbot fine-tuned on dialogue from Empresses in the Palace. huanhuan-chat stars
  • AMchat — An LLM-powered assistant for advanced mathematics. AMchat stars
  • d2l-ai-solutions-manual — Community solutions for Dive into Deep Learning exercises. d2l-ai-solutions-manual stars

Contributor

Research & Work Experience

Emotion Machine Lab — Researcher & Growth Operations

Present

  • Conduct research on Agent RL, LLM post-training, and agentic systems.
  • Work across developer growth, community operations, and the open-source ecosystem.

Westlake University, AGI Lab — Research Assistant

Yunqi Academy of Engineering — Research Assistant

Jun 2024 – Aug 2024

Selected Publications

GitHub Stats

KMnO4-zx's GitHub statsKMnO4-zx's most used languages

Pinned Loading

  1. datawhalechina/happy-llm datawhalechina/happy-llm Public

    📚 从零开始构建大模型

    Jupyter Notebook 32.4k 3.1k

  2. datawhalechina/self-llm datawhalechina/self-llm Public

    《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程

    Jupyter Notebook 31.4k 3.1k

  3. paper-insight paper-insight Public

    Paper Insight - AI驱动的学术论文智能分析

    TypeScript 128 8

  4. llm-agent-rl-lab llm-agent-rl-lab Public

    Reproducing and studying RL algorithms for LLM agents, including PPO, GRPO, GSPO, DAPO, OPD and beyond.

    Python 73 7