Skip to content

Repository files navigation

LLM Eval Tracker

Live index of LLM evaluation tools and benchmarks, refreshed every 15 minutes from GitHub

Stars Last Commit Items Updated

⭐ Star this repo to bookmark — fresh data every 15 minutes

English · 中文 · 日本語 · 한국어 · Español · Português


💡 What is this?

Automatically discovers and indexes new LLM evaluation frameworks, benchmarks, and harnesses as they appear on GitHub. Generates a structured, searchable catalog with metadata like stars, activity, and category tags. Designed for ML engineers who need to stay current without manually scanning repositories.

This list is auto-updated every 15 minutes by a GitHub Actions cron. Each commit reflects a real change in the upstream data source — new items added, expired items removed — so you can rely on what you see being current.


📋 Current Items

⏰ Last updated: 2026-09-01 04:30 UTC

Data source: GitHub Search API

The table below is rewritten on every cron tick. Star the repo to bookmark.

# Name Lang Updated Description
1 lokesh75-kank/agenteval 0 TypeScript 2026-09-01 Reliability and audit-evidence testing for LLM agents - wrap any agent, assert behavior, measure determinism, check grou
2 HaileyStorm/Creative-Writing-Rubrics 0 Python 2026-09-01 HBQ-RS: composable binary-question rubrics for creative writing, draft judging, benchmarking, and synthetic data.
3 gmitt98/fieldtest 0 Python 2026-09-01 LLM evaluation framework — define what correct, well-formed, and safe means before you measure
4 Arize-ai/phoenix 11267 Python 2026-09-01 AI Observability & Evaluation
5 promptfoo/promptfoo 24711 TypeScript 2026-09-01 Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C
6 Kondwani10/Origin-Continuum 0 2026-08-31 🌐 Define and explore the Origin ↔ Continuum framework, ensuring proper attribution and continuity in dependency relation
7 Steel-predictor-project/steel-llm-eval 0 JavaScript 2026-08-31 Open benchmark: how well can LLMs predict knife-steel properties (edge retention, toughness) from chemical composition,
8 truera/trulens 3531 Python 2026-08-31 Evaluation and Tracking for LLM Experiments and AI Agents
9 goldbarth/chartula-evals 0 Python 2026-08-31 How Chartula is measured: eval cases, run costs, and judgement of the generated changelogs.
10 homayoun-safarpour/homayoun-safarpour 0 2026-08-31 judge-drift-sentinel · judge-reliability-kit · agent-loop-engine · trace-gate · ai-eng-skill-range
11 Giskard-AI/giskard-oss 5799 Python 2026-08-31 🐢 Open-Source Evaluation & Testing library for LLM Agents
12 pdxlab/trustmodel-mcp-server 0 TypeScript 2026-08-31 TrustModel MCP Server — trust evaluation, red-team, and governance for AI agents via the Model Context Protocol. npm: @t
13 cannonade-ai/cannonade 2 TypeScript 2026-08-31 Local-first desktop app for building LLM test suites and running them against many local or cloud models at once
14 verifywise-ai/verifywise 343 TypeScript 2026-08-31 Complete AI governance and LLM Evals platform with support for EU AI Act, ISO 42001, NIST AI RMF and 20+ more AI framewo
15 decimal-labs/decimalai-python 1 Python 2026-08-31 📐 Python SDK for agent evals and skill routing — measure a skill's real lift before you trust it
16 ahmedmoha9088/PhoenixFish 0 Java 2026-08-31 Elevate your Paper server with an immersive fishing overhaul featuring custom fish, rods, bait, and a dynamic minigame.
17 saddled-panicattack529/idea-evaluation-pipeline 0 2026-08-31 Streamline research idea evaluation for finance and economics to reach top journal quality using an iterative, AI-assist
18 SFX-TECH/sfx-lead-intelligence 0 2026-08-31 SFX Lead Intelligence Command Center: local-LLM hub plus lead dashboard, quality lifted 61 to 99 percent via a ground-tr
19 lihongyu-dev/anzhice-crm 0 TypeScript 2026-08-30 贷款线索 CRM:LLM 资质抽取 + eval 框架 + 规则引擎撮合判定
20 Sans-cell-art/-Project-Phoenix-The-E-Waste-Supercomputer- 0 2026-08-30 ♻️ Transform e-waste into a powerful, low-cost cloud operating system, unlocking computing potential and promoting resou
21 bhavya7995/AI_governance 1 PowerShell 2026-08-30 🤖 Streamline AI-assisted development with a governance kit for rules, enforcement, and decision-making, ensuring speed a
22 Benj124/judge-dredd 0 TypeScript 2026-08-29 Local kit for ingesting a corpus, synthesizing LLM eval questions, reviewing gold, and judging model answers.
23 ChelseaKR/plumbline 1 Python 2026-08-29 v0.2.0. Fail-closed evaluation harness for government-facing chat systems: reproducible, provenance-stamped audit verdic
24 ChelseaKR/sprout 1 Python 2026-08-29 In-build reference implementation: an offline-first plant-care assistant and public evaluation harness with cited-corpus
25 ChelseaKR/gauntlet 1 Python 2026-08-29 v0.1.0. Merge-blocking evaluation gates for generative AI features: YAML suites run against any HTTP endpoint or Python
26 ChelseaKR/fare-policy-assistant 1 HTML 2026-08-29 Beta. Bilingual reduced-fare policy assistant grounded in dated citations, with a corpus of eighteen California transit
27 izam-mohammed/ragrank 47 Python 2026-08-29 🎯 Your free LLM evaluation toolkit helps you assess the accuracy of facts, how well it understands context, its tone, an
28 camerontjs-dot/agent-eval-notes 0 CSS 2026-08-28 Public-safe agent evaluation write-ups: harness gates, multi-path coding screens, task-family transfer, RAG routes, agen
29 IonDen/mlx-quant-fidelity 4 Python 2026-08-28 Measure quantization quality loss on Apple Silicon MLX — KL divergence, top-token flip rate and perplexity delta for KV-
30 lordbasilaiassistant-sudo/company-bench 3 JavaScript 2026-08-28 Can your AI agent hold a job? Open-source benchmark for AI agent trustworthiness, not capability: 29 chairs, 7 departmen
31 vishalmurugan1986/support-tam-ai-copilot 1 Python 2026-08-27 LLM-powered ticket triage + TAM account health briefs for technical support teams, with a rule-based fallback, RAG over
32 jafeeri/llm-eval-bench 0 Python 2026-08-27 Score LLM outputs and block quality regressions in CI. Deterministic checks first, a calibrated LLM judge where needed.
33 camerontjs-dot/verified-done 0 Python 2026-08-28 Does done mean done? Coding-agent honesty demo: verified pass vs false completion vs scope violation. Public demo split
34 isatimur/mash-core 0 Python 2026-08-25
35 jeremylongshore/j-rig-skill-binary-eval 1 TypeScript 2026-08-31 Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score e
36 RudrenduPaul/memtrust 1 Python 2026-08-25 Independent CLI benchmark harness for agent-memory backends (MemPalace, Mem0, Zep, OpenViking); publishes raw eval logs.
37 lorien/awesome-ai-benchmarks 1 2026-08-25 Curated list of benchmarks and rankings of models, agents and other AI-things.
38 sunxin-ai/dsh-design-qa 3 JavaScript 2026-08-25 Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches
39 rithvik-bk/voiceos-eval 0 JavaScript 2026-08-25 Tool-calling eval harness + pre-execution safety gate for voice agents: scores whether the right tool was called with th
40 isatimur/book-mash 0 Python 2026-08-22
41 ozlar34/job-match-radar 1 Python 2026-08-22 Self-hosted n8n + Supabase pipeline that scrapes LinkedIn and a watchlist of company ATS endpoints, scores listings agai
42 fabio-barboza/logistic-platform 0 Java 2026-08-22 Agente de IA para logistica: chat em linguagem natural sobre frota, rotas e entregas. Java 21, Spring Boot 4, Spring Sec
43 tainguyen07/llm-eval-harness 4 Python 2026-08-21 Evaluation harness for LLM apps: prompt versioning, dataset runners, LLM-as-judge scoring, regression tracking, and HTML
44 multivon-ai/multivon-eval 25 Python 2026-08-20 Practical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI
45 eliasfeitan-pixel/llm-eval-framework 1 2026-08-20 Production-grade evaluation framework and automated guardrails for enterprise LLM applications, RAG pipelines, and agent
46 Player-YN/TokSight 0 Python 2026-08-19 Local Outcome lab for agent workflows. Import yours, Run N, judge with validate(output) or human Pass/Fail, export trace
47 sx4im/skillcheck 26 TypeScript 2026-08-25 A/B test agent skills with blind grading + bootstrap CIs - does your SKILL.md actually improve task performance?
48 j-newcom/retail-cpg-eval-datasets 0 Python 2026-08-18 Open, domain-specific evaluation datasets and binary judges for Retail & CPG generative-AI tasks.
49 modwin/agent-eval-framework 0 Python 2026-08-18 LLM Agent Evaluation Platform
50 gabriel-ngrs/CalorIA 1 Python 2026-08-24 Diário alimentar com IA: eval do pipeline de LLM versionado junto do código

🔍 How it works

Every 15 minutes, a GitHub Action runs tracker.py. That script:

  1. Fetches the latest state from GitHub Search API.
  2. Diffs against data/items.json (the previous snapshot).
  3. Rewrites the table above between the <!-- TRACKER_TABLE_* --> markers.
  4. Commits feat: +N added, -M removed (timestamp) if anything changed.

No external services. No paid APIs. Just a public data source and a free GitHub Action.


🤝 Contributing

See CONTRIBUTING.md — usually you don't need to: the tracker keeps itself current. If you spot a data-source bug or want to suggest a new column for the table, open an issue.


🔗 Related live trackers

If you find this useful, you might also like these other auto-updated trackers from the same maintainer — same mechanism, different upstream:


📜 License

MIT — see LICENSE.