Skip to content
#

agent-evaluation

Here are 1,113 public repositories matching this topic...

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

  • Updated Sep 18, 2026
  • Python

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

  • Updated Sep 22, 2026
  • Python
AgentMeasure

Open measurement infrastructure for AI agents. Our audit of 124 usage tools found 45+ verified billing bugs — 19 fixes landed upstream. Conformance fixtures for token accounting: PASS / FAIL / UNPROVABLE in CI.

  • Updated Sep 22, 2026
  • Python
coder_eval

Playwright for coding agents. Test that your skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, A/B experiments, CI gates.

  • Updated Sep 22, 2026
  • Python

Add this topic to your repo

To associate your repository with the agent-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more