Stars
SpaceXAI's coding agent harness and TUI. Fullscreen, mouse interactive, extensible.
Measuring frontier coding agents on original, long-horizon engineering tasks
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
All parts of Claude Code's system prompt, 27 builtin tool descriptions, sub agent prompts (Plan/Explore/Task), utility prompts (CLAUDE.md, compact, statusline, magic docs, WebFetch, Bash cmd, secur…
Harness for running and evaluating AI agents against RL environments
The Agentic Commerce Protocol (ACP) is an interaction model and open standard for connecting buyers, their AI agents, and businesses to complete purchases seamlessly. The specification is currently…
All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.
A collection of projects designed to help developers quickly get started with building deployable applications using the Claude API
Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.
AIDE: an LLM agent for machine learning engineering - the research Weco grew out of. Referenced in OpenAI MLE-bench.
Public repository containing METR's DVC pipeline for eval data analysis
Vivaria is METR's tool for running evaluations and conducting agent elicitation research.
[NeurIPS'25] Official codebase for "SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution"
Collection of leaked system prompts
Build PowerPoint presentations with JavaScript. Works with Node, React, web browsers, and more.
Lightweight coding agent that runs in your terminal
🌐 Make websites accessible for AI agents. Automate tasks online with ease.
Tool for data extraction and interacting with Lean programmatically.