Live index of LLM evaluation tools and benchmarks, refreshed every 15 minutes from GitHub
⭐ Star this repo to bookmark — fresh data every 15 minutes
Automatically discovers and indexes new LLM evaluation frameworks, benchmarks, and harnesses as they appear on GitHub. Generates a structured, searchable catalog with metadata like stars, activity, and category tags. Designed for ML engineers who need to stay current without manually scanning repositories.
This list is auto-updated every 15 minutes by a GitHub Actions cron. Each commit reflects a real change in the upstream data source — new items added, expired items removed — so you can rely on what you see being current.
⏰ Last updated: 2026-09-01 04:30 UTC
Data source:
GitHub Search APIThe table below is rewritten on every cron tick. Star the repo to bookmark.
| # | Name | ⭐ | Lang | Updated | Description |
|---|---|---|---|---|---|
| 1 | lokesh75-kank/agenteval | 0 | TypeScript | 2026-09-01 | Reliability and audit-evidence testing for LLM agents - wrap any agent, assert behavior, measure determinism, check grou |
| 2 | HaileyStorm/Creative-Writing-Rubrics | 0 | Python | 2026-09-01 | HBQ-RS: composable binary-question rubrics for creative writing, draft judging, benchmarking, and synthetic data. |
| 3 | gmitt98/fieldtest | 0 | Python | 2026-09-01 | LLM evaluation framework — define what correct, well-formed, and safe means before you measure |
| 4 | Arize-ai/phoenix | 11267 | Python | 2026-09-01 | AI Observability & Evaluation |
| 5 | promptfoo/promptfoo | 24711 | TypeScript | 2026-09-01 | Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C |
| 6 | Kondwani10/Origin-Continuum | 0 | — | 2026-08-31 | 🌐 Define and explore the Origin ↔ Continuum framework, ensuring proper attribution and continuity in dependency relation |
| 7 | Steel-predictor-project/steel-llm-eval | 0 | JavaScript | 2026-08-31 | Open benchmark: how well can LLMs predict knife-steel properties (edge retention, toughness) from chemical composition, |
| 8 | truera/trulens | 3531 | Python | 2026-08-31 | Evaluation and Tracking for LLM Experiments and AI Agents |
| 9 | goldbarth/chartula-evals | 0 | Python | 2026-08-31 | How Chartula is measured: eval cases, run costs, and judgement of the generated changelogs. |
| 10 | homayoun-safarpour/homayoun-safarpour | 0 | — | 2026-08-31 | judge-drift-sentinel · judge-reliability-kit · agent-loop-engine · trace-gate · ai-eng-skill-range |
| 11 | Giskard-AI/giskard-oss | 5799 | Python | 2026-08-31 | 🐢 Open-Source Evaluation & Testing library for LLM Agents |
| 12 | pdxlab/trustmodel-mcp-server | 0 | TypeScript | 2026-08-31 | TrustModel MCP Server — trust evaluation, red-team, and governance for AI agents via the Model Context Protocol. npm: @t |
| 13 | cannonade-ai/cannonade | 2 | TypeScript | 2026-08-31 | Local-first desktop app for building LLM test suites and running them against many local or cloud models at once |
| 14 | verifywise-ai/verifywise | 343 | TypeScript | 2026-08-31 | Complete AI governance and LLM Evals platform with support for EU AI Act, ISO 42001, NIST AI RMF and 20+ more AI framewo |
| 15 | decimal-labs/decimalai-python | 1 | Python | 2026-08-31 | 📐 Python SDK for agent evals and skill routing — measure a skill's real lift before you trust it |
| 16 | ahmedmoha9088/PhoenixFish | 0 | Java | 2026-08-31 | Elevate your Paper server with an immersive fishing overhaul featuring custom fish, rods, bait, and a dynamic minigame. |
| 17 | saddled-panicattack529/idea-evaluation-pipeline | 0 | — | 2026-08-31 | Streamline research idea evaluation for finance and economics to reach top journal quality using an iterative, AI-assist |
| 18 | SFX-TECH/sfx-lead-intelligence | 0 | — | 2026-08-31 | SFX Lead Intelligence Command Center: local-LLM hub plus lead dashboard, quality lifted 61 to 99 percent via a ground-tr |
| 19 | lihongyu-dev/anzhice-crm | 0 | TypeScript | 2026-08-30 | 贷款线索 CRM:LLM 资质抽取 + eval 框架 + 规则引擎撮合判定 |
| 20 | Sans-cell-art/-Project-Phoenix-The-E-Waste-Supercomputer- | 0 | — | 2026-08-30 | ♻️ Transform e-waste into a powerful, low-cost cloud operating system, unlocking computing potential and promoting resou |
| 21 | bhavya7995/AI_governance | 1 | PowerShell | 2026-08-30 | 🤖 Streamline AI-assisted development with a governance kit for rules, enforcement, and decision-making, ensuring speed a |
| 22 | Benj124/judge-dredd | 0 | TypeScript | 2026-08-29 | Local kit for ingesting a corpus, synthesizing LLM eval questions, reviewing gold, and judging model answers. |
| 23 | ChelseaKR/plumbline | 1 | Python | 2026-08-29 | v0.2.0. Fail-closed evaluation harness for government-facing chat systems: reproducible, provenance-stamped audit verdic |
| 24 | ChelseaKR/sprout | 1 | Python | 2026-08-29 | In-build reference implementation: an offline-first plant-care assistant and public evaluation harness with cited-corpus |
| 25 | ChelseaKR/gauntlet | 1 | Python | 2026-08-29 | v0.1.0. Merge-blocking evaluation gates for generative AI features: YAML suites run against any HTTP endpoint or Python |
| 26 | ChelseaKR/fare-policy-assistant | 1 | HTML | 2026-08-29 | Beta. Bilingual reduced-fare policy assistant grounded in dated citations, with a corpus of eighteen California transit |
| 27 | izam-mohammed/ragrank | 47 | Python | 2026-08-29 | 🎯 Your free LLM evaluation toolkit helps you assess the accuracy of facts, how well it understands context, its tone, an |
| 28 | camerontjs-dot/agent-eval-notes | 0 | CSS | 2026-08-28 | Public-safe agent evaluation write-ups: harness gates, multi-path coding screens, task-family transfer, RAG routes, agen |
| 29 | IonDen/mlx-quant-fidelity | 4 | Python | 2026-08-28 | Measure quantization quality loss on Apple Silicon MLX — KL divergence, top-token flip rate and perplexity delta for KV- |
| 30 | lordbasilaiassistant-sudo/company-bench | 3 | JavaScript | 2026-08-28 | Can your AI agent hold a job? Open-source benchmark for AI agent trustworthiness, not capability: 29 chairs, 7 departmen |
| 31 | vishalmurugan1986/support-tam-ai-copilot | 1 | Python | 2026-08-27 | LLM-powered ticket triage + TAM account health briefs for technical support teams, with a rule-based fallback, RAG over |
| 32 | jafeeri/llm-eval-bench | 0 | Python | 2026-08-27 | Score LLM outputs and block quality regressions in CI. Deterministic checks first, a calibrated LLM judge where needed. |
| 33 | camerontjs-dot/verified-done | 0 | Python | 2026-08-28 | Does done mean done? Coding-agent honesty demo: verified pass vs false completion vs scope violation. Public demo split |
| 34 | isatimur/mash-core | 0 | Python | 2026-08-25 | |
| 35 | jeremylongshore/j-rig-skill-binary-eval | 1 | TypeScript | 2026-08-31 | Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score e |
| 36 | RudrenduPaul/memtrust | 1 | Python | 2026-08-25 | Independent CLI benchmark harness for agent-memory backends (MemPalace, Mem0, Zep, OpenViking); publishes raw eval logs. |
| 37 | lorien/awesome-ai-benchmarks | 1 | — | 2026-08-25 | Curated list of benchmarks and rankings of models, agents and other AI-things. |
| 38 | sunxin-ai/dsh-design-qa | 3 | JavaScript | 2026-08-25 | Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches |
| 39 | rithvik-bk/voiceos-eval | 0 | JavaScript | 2026-08-25 | Tool-calling eval harness + pre-execution safety gate for voice agents: scores whether the right tool was called with th |
| 40 | isatimur/book-mash | 0 | Python | 2026-08-22 | |
| 41 | ozlar34/job-match-radar | 1 | Python | 2026-08-22 | Self-hosted n8n + Supabase pipeline that scrapes LinkedIn and a watchlist of company ATS endpoints, scores listings agai |
| 42 | fabio-barboza/logistic-platform | 0 | Java | 2026-08-22 | Agente de IA para logistica: chat em linguagem natural sobre frota, rotas e entregas. Java 21, Spring Boot 4, Spring Sec |
| 43 | tainguyen07/llm-eval-harness | 4 | Python | 2026-08-21 | Evaluation harness for LLM apps: prompt versioning, dataset runners, LLM-as-judge scoring, regression tracking, and HTML |
| 44 | multivon-ai/multivon-eval | 25 | Python | 2026-08-20 | Practical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI |
| 45 | eliasfeitan-pixel/llm-eval-framework | 1 | — | 2026-08-20 | Production-grade evaluation framework and automated guardrails for enterprise LLM applications, RAG pipelines, and agent |
| 46 | Player-YN/TokSight | 0 | Python | 2026-08-19 | Local Outcome lab for agent workflows. Import yours, Run N, judge with validate(output) or human Pass/Fail, export trace |
| 47 | sx4im/skillcheck | 26 | TypeScript | 2026-08-25 | A/B test agent skills with blind grading + bootstrap CIs - does your SKILL.md actually improve task performance? |
| 48 | j-newcom/retail-cpg-eval-datasets | 0 | Python | 2026-08-18 | Open, domain-specific evaluation datasets and binary judges for Retail & CPG generative-AI tasks. |
| 49 | modwin/agent-eval-framework | 0 | Python | 2026-08-18 | LLM Agent Evaluation Platform |
| 50 | gabriel-ngrs/CalorIA | 1 | Python | 2026-08-24 | Diário alimentar com IA: eval do pipeline de LLM versionado junto do código |
Every 15 minutes, a GitHub Action runs tracker.py. That script:
- Fetches the latest state from
GitHub Search API. - Diffs against
data/items.json(the previous snapshot). - Rewrites the table above between the
<!-- TRACKER_TABLE_* -->markers. - Commits
feat: +N added, -M removed (timestamp)if anything changed.
No external services. No paid APIs. Just a public data source and a free GitHub Action.
See CONTRIBUTING.md — usually you don't need to: the tracker keeps itself current.
If you spot a data-source bug or want to suggest a new column for the table, open
an issue.
If you find this useful, you might also like these other auto-updated trackers from the same maintainer — same mechanism, different upstream:
- trending-claude-skills — What's shipping in Claude Skills this week (
topic:claude-skills) - mcp-servers-live — Live index of newest MCP servers (
topic:mcp-server) - cursor-rules-live — Newest Cursor rules and .cursorrules patterns (
topic:cursor-rules) - claude-code-plugin-tracker — Claude Code plugins and hook configs (
topic:claude-code) - llm-agents-radar — Newest LLM agent frameworks (
topic:llm-agent) - rag-radar — Newest RAG implementations and tools (
topic:rag) - agent-framework-radar — Newest agent frameworks shipping on GitHub (
topic:agent-framework) - vector-db-live — Newest vector DB projects and integrations (
topic:vector-database) - llmops-radar — Newest LLMOps tooling (observability, deployment) (
topic:llmops) - prompt-tools-live — Newest prompt-engineering tools and prompt repos (
topic:prompt-engineering) - agent-eval-harness — Live benchmark of AI coding agents (
topic:llm-eval) - skills-tracker — Tracking new GitHub 'skills' repos (
topic:agent-skills) - awesome-agent-skills — Curated auto-updated awesome-list of AI agent skills (
topic:agent-skills)
MIT — see LICENSE.