Stars
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.
Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executa…
Measuring and evolving with the frontier of agent work
🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team.
Samaya AI's FrontierFinance Benchmark Grader
ReactBench is an evaluation for coding agents on realistic React work
A Kotlin software-engineering benchmark for evaluating coding agents.
Code for paper AIDABench: AI Data Analytics Benchmark.
WideSearch: Benchmarking Agentic Broad Info-Seeking
Terminal-Bench Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal
Real-world buy-side tasks for evaluating agents on investment research. China-centered, globally aware (A-share · HK · US · macro). Public set.
A benchmark for evaluating AI agents on realistic business workflows
The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a single command, no API glue code.
Framework for evaluating and improving agents
SWE-Marathon: an ultra long-horizon SWE benchmark
Mercor-Intelligence / harbor
Forked from harbor-framework/harborHarbor is a framework for running agent evaluations and creating and using RL environments.
Harness for running and evaluating AI agents against RL environments
Benchmark Everything Everywhere All at Once, a fully autonomous agentic system for benchmark construction and customization.
Anonymous Github is a proxy server to support anonymous browsing of Github repositories for open-science code and data.
Goal-driven harness experiments on real agent runtimes — a local, stdlib-only CLI workbench with Copilot + Auto modes and evidence-first reporting.
Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consistent agent trajectories
Measuring frontier coding agents on original, long-horizon engineering tasks
A coding agent framework, that works on its own codebase.
Benchmarking LLMs on Real-World Software Verification in Lean 4
Lightweight coding agent that runs in your terminal