Highlights
- Pro
Stars
DeepSeek Harness: Everything is a Plugin.
Diving into Reliable Self-Evolving Agents: A Survey
Official AHE code — Agentic Harness Engineering: observability-driven automatic evolution of coding-agent harnesses (concurrent w/ meta-harness). NexAU-AHE reaches 84.7% ± 2.1 pass@1 on Terminal-Be…
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
A benchmark of real-world DL kernel problems
A self-evolving multi-agent system for autonomous research, operating 24/7 to explore, learn, and improve.
Agentic RL on Any Harness at Scale
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
The roadmap of long-horizon agents
Synthetic data generation, post-training, and E2B benchmark evaluation infrastructure.
🗓️ The hardest life-admin benchmark for agents — lawsuits, escrow shortfalls, apartment hunts, exams. 20 long-horizon tasks × 20–30 stages across 10 domains and 21 services, scored by 1247 atomic c…
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
[ICLR 2026] VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
Ď„-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Official implementation of CORE (ICLR 2026) — Concept-Oriented Reinforcement for bridging the definition–application gap in mathematical reasoning. Concept-guided GRPO (CORE-Base/CR/KL), SC@21 eval…
Post-training with Tinker
The agent that grows with you
AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and supp…
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
AgentCPM-GUI: An on-device GUI agent for operating Android apps, enhancing reasoning ability with reinforcement fine-tuning for efficient task execution.
My learning notes for ML SYS.
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
from vibe coding to agentic engineering - practice makes claude perfect
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepowe…
Code and implementations for the paper "AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning" by Zhiheng Xi et al.