🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
-
Updated
Sep 20, 2026 - TypeScript
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
AI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
Connect Your Agents And Harnesses With Any Provider 🦚
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.
Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.
Open-source observability for AI agents. Find where your agents fail, dispatch your coding agent to fix it, and verify the fix against real traces.
AI observability platform for production LLM and agent systems.
Agent Skills as a Memory Layer
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
Open-source observability & evaluation platform for AI agents and coding agents. Trace LLMs, tools, prompts, costs & agent workflows with OpenTelemetry.
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Enterprise AI bastion host for secure AI API and MCP access, with unified proxying, RBAC, audit logs, rate limiting, and cost tracking across OpenAI, Anthropic, Gemini, and self-hosted LLMs.
TraceRoot - open-source observability and self-improving layer for AI agents. YC S25
🦞 Official plugin for OpenClaw that exports agent traces to Opik. See and monitor agent behaviour, cost, tokens, errors and more.
DataBuff is an AI-native APM built on Opentelemetry,with multi-agent troubleshooting out of the box.
To associate your repository with the llm-observability topic, visit your repo's landing page and select "manage topics."