Stars
DeepSeek Harness: Everything is a Plugin.
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.
Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executa…
Measuring and evolving with the frontier of agent work
🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team.
Community plugin to control Blender 3D with any LLM of your choice
Samaya AI's FrontierFinance Benchmark Grader
ReactBench is an evaluation for coding agents on realistic React work
A Kotlin software-engineering benchmark for evaluating coding agents.
Code for paper AIDABench: AI Data Analytics Benchmark.
WideSearch: Benchmarking Agentic Broad Info-Seeking
Terminal-Bench-Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal
Real-world buy-side tasks for evaluating agents on investment research. China-centered, globally aware (A-share · HK · US · macro). Public set.
A benchmark for evaluating AI agents on realistic business workflows
The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a single command, no API glue code.
Framework for evaluating and improving agents
SWE-Marathon: an ultra long-horizon SWE benchmark
Mercor-Intelligence / harbor
Forked from harbor-framework/harborHarbor is a framework for running agent evaluations and creating and using RL environments.
Harness for running and evaluating AI agents against RL environments
Benchmark Everything Everywhere All at Once, a fully autonomous agentic system for benchmark construction and customization.
Anonymous Github is a proxy server to support anonymous browsing of Github repositories for open-science code and data.
Goal-driven harness experiments on real agent runtimes — a local, stdlib-only CLI workbench with Copilot + Auto modes and evidence-first reporting.
Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consistent agent trajectories