- Beijing
Lists (1)
Sort Name ascending (A-Z)
Stars
The lightweight framework for building agents
Cutting-edge platform for LLM agent tuning. Deliver RL tuning with flexibility, reliability, speed, multi-agent optimization and realtime community benchmarking.
Scalable Agentic RL for Any Agent and Sandbox.
An LLM post-training framework with vLLM for RL Scaling
Agentic RL on Any Harness at Scale
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
The agent that grows with you
本地优先的跨平台 Claude Code / Agent 桌面工作台:多 Agent、Git Worktree、代码 Diff、技能市场、多模型、Computer Use、任务感知桌面宠物,并支持微信、飞书、钉钉、Telegram、WhatsApp 与 H5 访问。
OpenClaw-RL: Train any agent simply by talking
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
SkyRL: A Modular Full-stack RL Library for LLMs
[ICLR 2026] Tree Search for LLM Agent Reinforcement Learning
Codes for the paper "BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping" by Zhiheng Xi et al.
Tongyi Deep Research, the Leading Open-source Deep Research Agent
[ICLR 2026] Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
[ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
Bridge Megatron-Core to Hugging Face/Reinforcement Learning
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings
slime is an LLM post-training framework for RL Scaling.
A high-throughput and memory-efficient inference and serving engine for LLMs
Scalable toolkit for efficient model reinforcement
Muon is an optimizer for hidden layers in neural networks
Unleashing the Power of Reinforcement Learning for Math and Code Reasoners
Parallel Scaling Law for Language Model — Beyond Parameter and Inference Time Scaling
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.