Skip to content
View gongel's full-sized avatar
🎯
Focusing
🎯
Focusing

Organizations

@PaddlePaddle

Block or report gongel

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The lightweight framework for building agents

Python 525 58 Updated Jul 27, 2026

Cutting-edge platform for LLM agent tuning. Deliver RL tuning with flexibility, reliability, speed, multi-agent optimization and realtime community benchmarking.

Python 230 26 Updated Jul 2, 2026

Scalable Agentic RL for Any Agent and Sandbox.

Python 468 47 Updated Jul 28, 2026

An LLM post-training framework with vLLM for RL Scaling

Python 390 69 Updated Jul 28, 2026

Agentic RL on Any Harness at Scale

Python 714 75 Updated Jul 15, 2026

An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Python 546 121 Updated Jul 28, 2026

The agent that grows with you

Python 221,740 42,400 Updated Jul 28, 2026

本地优先的跨平台 Claude Code / Agent 桌面工作台:多 Agent、Git Worktree、代码 Diff、技能市场、多模型、Computer Use、任务感知桌面宠物,并支持微信、飞书、钉钉、Telegram、WhatsApp 与 H5 访问。

TypeScript 13,668 8,490 Updated Jul 28, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,611 608 Updated May 23, 2026

The open source coding agent.

TypeScript 190,453 24,194 Updated Jul 28, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,642 759 Updated May 17, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,101 392 Updated Jul 28, 2026

[ICLR 2026] Tree Search for LLM Agent Reinforcement Learning

Python 390 39 Updated Jan 26, 2026

Codes for the paper "BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping" by Zhiheng Xi et al.

Python 94 6 Updated Jan 29, 2026

Tongyi Deep Research, the Leading Open-source Deep Research Agent

Python 19,746 1,502 Updated Feb 27, 2026

[ICLR 2026] Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents

Python 128 8 Updated Jul 14, 2026

[ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)

Python 1,093 57 Updated Jul 13, 2026

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

Python 511 52 Updated Mar 30, 2026

Bridge Megatron-Core to Hugging Face/Reinforcement Learning

Python 228 79 Updated Jun 15, 2026

Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings

TypeScript 10,419 954 Updated Jul 28, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,679 1,100 Updated Jul 24, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 87,449 19,971 Updated Jul 28, 2026
Python 62 5 Updated Jul 21, 2025

Scalable toolkit for efficient model reinforcement

Python 1,855 490 Updated Jul 28, 2026

Agentic RL Training at Scale

Python 1,757 375 Updated Jul 28, 2026

Muon is an optimizer for hidden layers in neural networks

Python 2,747 130 Updated May 24, 2026

Unleashing the Power of Reinforcement Learning for Math and Code Reasoners

Python 740 44 Updated Jun 6, 2025

Parallel Scaling Law for Language Model — Beyond Parameter and Inference Time Scaling

Python 480 26 Updated May 17, 2025

Democratizing Reinforcement Learning for LLMs

Python 5,740 596 Updated Jul 28, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,062 1,027 Updated Jul 28, 2026
Next