Public tools for LLM / agent evaluation and reliability. This GitHub profile page is only an index - clone a named project below, not this repo.
| Project | Job |
|---|---|
| judge-drift-sentinel | Score moved: system or LLM judge? |
| judge-reliability-kit | Why a judge panel disagrees (kappa) |
| agent-loop-engine | State, gates, decide, journal |
| agent-loop-field-guide | Loop Contract before you automate |
| trace-gate | Trajectory deploy gate (exit 0/2) |
| rag-eval-service | RAG path + frozen hit@k / MRR |
| agent-eval-workbench | Scenario traces + detectors |
| repro-ml-pipeline | Train, register, serve + signature |
| ai-eng-skill-range | 56 skills / 24 graded katas |
| agent-constraint-auditor | Audit agent transcripts for declared-constraint decay |
| judge-field-guide | Link-checked map of the judge tool ecosystem |
Python 3.10-3.12 CI. Named tests behind README claims. Start with any Quickstart (under 30 minutes).