Long-horizon agent control plane for durable, governed work across Codex, Claude Code, and other harnesses.
-
Updated
Sep 21, 2026 - Python
Long-horizon agent control plane for durable, governed work across Codex, Claude Code, and other harnesses.
The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fresh-context execution, durable verified state, independent auditing, recoverable progress, and native Claude Code / Codex / OpenClaw integration.
Long Horizon Terminal Benchmark with Dense Reward Grading
High-performance AI agent for long-horizon tasks. Built on empirical research. 200k+ tokens of work inside a 64k context window.
The open-source design agent and harness, better than Claude Design on academic communication artifacts production. This DesignHarness can also be used with any coding harness you like ( Codex/Claude Code/Kimi Code/Pi/OpenCode etc..) and any agentic model you want.
Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses
AI4AI Survey: can AI reliably improve AI? 223 papers on long-horizon agents, benchmarks, harness design, and recursive self-improvement · updated weekly
The local-first intelligence layer that gives AI agents durable continuity, explainable retrieval, portable context, and verified learning across harnesses.
SpineCodex: Let your Codex work, evolve, and scale on a SpineTree — up to 10× effective context and 89% more SWE-Milestone tasks resolved at 27% lower cost.
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
Cayu is the runtime for long-horizon agents that need explicit environments, durable sessions, controlled tools, secrets boundaries, evals, and replay
[EMNLP 2026] LightRSI is a modular runtime for lightly deploying recursive self-improvement loops in long-horizon LLM agents.
EvoX Genesis is an autonomous system for long-horizon software evolution that recursively builds, continues, and transforms complex software from high-level objectives.
🔁 Build reliable recurring AI-agent systems: 1022 resources, 22 operational patterns, 22 loop contracts, 8 runtime starters, an interactive atlas, and a structured dataset.
Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Simple Long Horizon Agent - A simple yet effective AI agent for learning, experimentation, and long horizon work.
A workflow specification for autonomous agents
Dev skills I use day to day in agentic engineering, focused on multi-agent-driven work modes and long-horizon autonomous tasks.
Official code for Behavior-Skill, a fine-grained skill dataset and evaluation benchmark for Vision-Language-Action policies in long-horizon mobile manipulation tasks.
Benchmark-as-Teacher: self-evolving post-training for medical agents
To associate your repository with the long-horizon-agents topic, visit your repo's landing page and select "manage topics."