Skip to content

Trust score

The measured accuracy of Cogeto, per release

Current release v1.1.02026-07-25

Cogeto publishes its own measured accuracy for every release, the same way a service publishes uptime. Here are the numbers, and here are the public data files behind them. Do not trust this chart: check the file.

Model configuration
Language

Aggregate blends the per-language corpora. It is shown so a weak language can never hide inside an average: switch the selector to read each language on its own.

Current scores

Extraction and reconciliation quality for the selected model configuration and language, measured against a hand-labeled golden corpus.

Current scores for configuration mistral-default, Aggregate, release v1.1.0
MetricAggregate
78.9%
93.3%
86.5%
92.9%
100%

Chat suite

27/27cases pass

End-to-end question-and-answer cases. A pass means the answer was grounded in the right remembered facts. Failing case ids are published, not hidden.

Trends

Every published release, oldest to newest, on an honest 0 to 100 percent axis. The dashed line is the continuous-integration gate that a release must clear to ship.

Measured at releaseBackfilled: transcribed from recorded runs rather than emitted by the harness at release time.CI gate: minimum to ship
Extraction precision78.9%
Extraction recall93.3%
Verification agreement86.5%
Deduplication accuracy92.9%
Contradiction recall100%
Trend data for configuration mistral-default, Aggregate, all releases
ReleaseExtraction precisionExtraction recallVerification agreementDeduplication accuracyContradiction recall
v0.8.0 (Backfilled)77.8%89.7%95.5%92.9%100%
v0.9.1 (Backfilled)81%87.8%90%92.9%100%
v0.9.280%89.2%88%92.9%100%
v1.0.084%90%83.3%92.9%83.3%
v1.0.182.5%91.1%89.4%92.9%100%
v1.0.281.7%92.2%87.9%92.9%100%
v1.0.382.9%93.3%83.3%92.9%100%
v1.0.482.9%93.3%86.4%92.9%100%
v1.0.582.7%92.2%84.8%92.9%100%
v1.1.078.9%93.3%86.5%92.9%100%

Notes from the releases

  • v0.9.2Open the JSON file
    • Published via the manual retry path: the release-time emission hit a Mistral 429 when the tag push ran the live gate and the trust job concurrently against one API key (fixed by the shared live-model-eval concurrency group in this same commit).
    • Chat 11/12: atlas_scope graded 67% coverage on this run (83-100% in adjacent runs), coverage-grader variance on a small-fact case, not a behavior change; the failing case id is published per the honesty rule.
  • v0.9.1BackfilledOpen the JSON file
    • Backfilled: transcribed from the live CI gate run on main after the v0.9.1-era reply-intent fix (#79), the first fully green live-gate run.
    • Verification agreement dipped vs v0.8.0 (95.5% → 90.0%): the corpus grew from 46 to 52 cases with harder email-sourced Croatian cases (O4); the gate floor (75%) holds with margin.
    • v0.9.0 has no trust-score file: it shipped before the live gate had an API key on this repository, so no measured run exists for that tag.
  • v0.8.0BackfilledOpen the JSON file
    • Backfilled: golden-set and reconciliation numbers transcribed from the 2026-07-13 measured run in docs/eval/history.md (the run closest to the v0.8.0 tag).
    • Chat summary transcribed from the nearest measured chat run (2026-07-10, 10 cases), the two email reply cases were added to the suite after v0.8.0.

Provenance

Each release, with the exact commit it was measured at, the harness version, the corpus sizes, and a direct link to its immutable JSON file. Read the data, not our summary of it.

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0006 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 76 golden cases (37 english, 39 croatian) · 20 reconciliation pairs · 27 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 68 golden cases (33 english, 35 croatian) · 20 reconciliation pairs · 12 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0004 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 52 golden cases (32 english, 20 croatian) · 20 reconciliation pairs · 12 chat cases

  • v0.9.12026-07-15BackfilledOpen the JSON file5b6bdb3208

    Transcribed from recorded runs rather than emitted by the harness at release time.

    Harness
    transcribed: extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 · chat answer/v0004 + eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 52 golden cases (32 english, 20 croatian) · 20 reconciliation pairs · 12 chat cases

  • v0.8.02026-07-14BackfilledOpen the JSON file18c4f1c336

    Transcribed from recorded runs rather than emitted by the harness at release time.

    Harness
    transcribed: extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 46 golden cases (29 english, 17 croatian) · 18 reconciliation pairs · 10 chat cases

Back to cogeto.eu