Stars
Evaluation harness for Apodex-1.0 on public deep-research benchmarks.
large language model internal-medicine monitor toolbox
agentUniverse is a LLM multi-agent framework that allows developers to easily build multi-agent applications.
Structured deep research skill for Claude Code/Open Code/Codex with human-in-the-loop control
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
Framework for evaluating and improving agents
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Code and implementations for the paper "AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning" by Zhiheng Xi et al.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
slime is an LLM post-training framework for RL Scaling.
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
SGLang is a high-performance serving framework for large language models and multimodal models.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
An elegant \LaTeX\ résumé template. 大陆镜像 https://gods.coding.net/p/resume/git
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."
High-performance safetensors model loader
DeepEP: an efficient expert-parallel communication library
📄 适合中文的简历模板收集(LaTeX,HTML/JS and so on)由 @hoochanlon 维护
A bidirectional pipeline parallelism algorithm for computation-communication overlap in DeepSeek V3/R1 training.