Framework for evaluating and improving agents
-
Updated
Sep 1, 2026 - Python
Framework for evaluating and improving agents
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
HuggingEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
A Universal Platform for Training and Evaluation of Mobile Interaction
A graphical interface for reinforcement learning and gym-based environments.
Interoperating between (Deep) Reiforcement Learning libraries
Gymnasium-style API standard for RL environment creation in JAX
Workspace manager for coding agents. Interactively solve and develop Harbor tasks.
Create new gridworld gym environments easily
Awesome Environment Scaling for AI Agents — a periodically updated survey and curated list of papers, projects, RL environments, world models, agent sandboxes, and open research problems
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
A configurable 3D jump-chess framework with A/B geometries, multi-player rule switches, and RL/self-play training support.
Agent-evaluation environments: planted-truth worlds, ungameable graders, calibrated difficulty
A lightweight, open-source framework that turns historical GitHub pull requests into reproducible, verifiable software-engineering tasks for training and evaluating coding agents.
Foundry Lite: a public runnable sample of Veyl’s local environment harness for software-engineering agent evals.
Open-source RL environments and evals for the capabilities we want AI to have. Judge-free scoring, mandatory baselines, and an enforced defensive-asymmetry gate.
Sound error bounds, symbolic GPU safety checks and targeted falsification for Triton kernels. 88% of planted bugs pass the standard fixed-shape allclose test; Litmus catches 100% with 0% false positives, on CPU.
An original Terminal-Bench 3 task that Claude Opus 5 and GPT-5.6 both fail three times out of three, and the research on why. The identity model lives in the data rather than the spec, and the graded core is opaque, so an agent cannot check its own answer.
To associate your repository with the rl-environments topic, visit your repo's landing page and select "manage topics."