AI Observability & Evaluation
-
Updated
Sep 20, 2026 - Python
AI Observability & Evaluation
OpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100+ datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.
🐢 Open-Source Evaluation & Testing library for LLM Agents
Evaluation and Tracking for LLM Experiments and AI Agents
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.
Python SDK for running evaluations on LLM generated responses
A simple GPT-based evaluation tool for multi-aspect, interpretable assessment of LLMs.
Python SDK for experimenting, testing, evaluating & monitoring LLM-powered applications - Parea AI (YC S23)
First-of-its-kind AI benchmark for evaluating the protection capabilities of large language model (LLM) guard systems (guardrails and safeguards)
llm-eval-simple is a simple LLM evaluation framework with intermediate actions and prompt pattern selection
🎯 Your free LLM evaluation toolkit helps you assess the accuracy of facts, how well it understands context, its tone, and more. This helps you see how good your LLM applications are.
Develop reliable AI apps
Valor is a lightweight, numpy-based library designed for fast and seamless evaluation of machine learning models.
An graph-eval framework for LLM's
An open source library for asynchronous querying of LLM endpoints
The warehouse-native LLM evaluation package for dbt™ - monitor AI quality without data egress
LLM Security Platform.
Structured output benchmarks comparing DSPy and BAML with different LLMs
SWE-bench for your codebase — mine your merged PRs into local, contamination-free coding-agent benchmarks. Adapters: claude-code, aider (Opus 4.7 / GPT-5.5 / Sonnet 4.6 / Gemini 3.1 Pro).
Practical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI/CD integration.
To associate your repository with the llm-eval topic, visit your repo's landing page and select "manage topics."