Skip to content

W4 — Eval harness + benchmarks

Open
No due date
Last updated Jul 26, 2026

Synthetic corpus generator. 10-scenario adversarial suite with naive RAG baseline. Four benchmark tables (overhead, recall vs selectivity, scaling, revocation propagation). make bench reproduces everything.

33% complete

List view