Trust score
The measured accuracy of Cogeto, per release
Cogeto publishes its own measured accuracy for every release, the same way a service publishes uptime. Here are the numbers, and here are the public data files behind them. Do not trust this chart: check the file.
Aggregate blends the per-language corpora. It is shown so a weak language can never hide inside an average: switch the selector to read each language on its own.
Current scores
Extraction and reconciliation quality for the selected model configuration and language, measured against a hand-labeled golden corpus.
| Metric | Aggregate | CI gate |
|---|---|---|
| 78.9% | CI gate ≥ 70% (clears the gate) | |
| 93.3% | CI gate ≥ 80% (clears the gate) | |
| 86.5% | CI gate ≥ 75% (clears the gate) | |
| 92.9% | CI gate ≥ 90% (clears the gate) | |
| 100% | CI gate ≥ 70% (clears the gate) |
Chat suite
27/27cases pass
End-to-end question-and-answer cases. A pass means the answer was grounded in the right remembered facts. Failing case ids are published, not hidden.
Trends
Every published release, oldest to newest, on an honest 0 to 100 percent axis. The dashed line is the continuous-integration gate that a release must clear to ship.
| Release | Extraction precision | Extraction recall | Verification agreement | Deduplication accuracy | Contradiction recall |
|---|---|---|---|---|---|
| v0.8.0 (Backfilled) | 77.8% | 89.7% | 95.5% | 92.9% | 100% |
| v0.9.1 (Backfilled) | 81% | 87.8% | 90% | 92.9% | 100% |
| v0.9.2 | 80% | 89.2% | 88% | 92.9% | 100% |
| v1.0.0 | 84% | 90% | 83.3% | 92.9% | 83.3% |
| v1.0.1 | 82.5% | 91.1% | 89.4% | 92.9% | 100% |
| v1.0.2 | 81.7% | 92.2% | 87.9% | 92.9% | 100% |
| v1.0.3 | 82.9% | 93.3% | 83.3% | 92.9% | 100% |
| v1.0.4 | 82.9% | 93.3% | 86.4% | 92.9% | 100% |
| v1.0.5 | 82.7% | 92.2% | 84.8% | 92.9% | 100% |
| v1.1.0 | 78.9% | 93.3% | 86.5% | 92.9% | 100% |
Notes from the releases
- v0.9.2Open the JSON file
- Published via the manual retry path: the release-time emission hit a Mistral 429 when the tag push ran the live gate and the trust job concurrently against one API key (fixed by the shared live-model-eval concurrency group in this same commit).
- Chat 11/12: atlas_scope graded 67% coverage on this run (83-100% in adjacent runs), coverage-grader variance on a small-fact case, not a behavior change; the failing case id is published per the honesty rule.
- v0.9.1BackfilledOpen the JSON file
- Backfilled: transcribed from the live CI gate run on main after the v0.9.1-era reply-intent fix (#79), the first fully green live-gate run.
- Verification agreement dipped vs v0.8.0 (95.5% → 90.0%): the corpus grew from 46 to 52 cases with harder email-sourced Croatian cases (O4); the gate floor (75%) holds with margin.
- v0.9.0 has no trust-score file: it shipped before the live gate had an API key on this repository, so no measured run exists for that tag.
- v0.8.0BackfilledOpen the JSON file
- Backfilled: golden-set and reconciliation numbers transcribed from the 2026-07-13 measured run in docs/eval/history.md (the run closest to the v0.8.0 tag).
- Chat summary transcribed from the nearest measured chat run (2026-07-10, 10 cases), the two email reply cases were added to the suite after v0.8.0.
Provenance
Each release, with the exact commit it was measured at, the harness version, the corpus sizes, and a direct link to its immutable JSON file. Read the data, not our summary of it.
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0006 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 76 golden cases (37 english, 39 croatian) · 20 reconciliation pairs · 27 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 52 golden cases (32 english, 20 croatian) · 20 reconciliation pairs · 12 chat cases
Transcribed from recorded runs rather than emitted by the harness at release time.
- Harness
- transcribed: extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 · chat answer/v0004 + eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 52 golden cases (32 english, 20 croatian) · 20 reconciliation pairs · 12 chat cases
Transcribed from recorded runs rather than emitted by the harness at release time.
- Harness
- transcribed: extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 46 golden cases (29 english, 17 croatian) · 18 reconciliation pairs · 10 chat cases