Skip to content

Trust scores

The measured accuracy of Cogeto, per release

Current release v1.8.02026-08-16

Cogeto publishes its own measured accuracy for every release, the same way a service publishes uptime, including the numbers that fall short of their targets. Here are the numbers, and here are the public data files behind them. Do not trust this chart: check the file.

Model configuration
Language
Compare

Aggregate blends the per-language corpora. It is shown so a weak language can never hide inside an average: switch the selector to read each language on its own.

Every gate floor is set at the honest current value of the metric, never at a target the project has not reached, and floors only ratchet upward. Floors apply per language as well as in aggregate, so the gate you see here changes with the language you select.

Current scores

Extraction and reconciliation quality for the selected model configuration and language, measured against a hand-labeled golden corpus.

Scores for configuration mistral-default, Aggregate, v1.8.0
MetricAggregate
84.1%
94.6%
91.9%
94.4%
81.3%
92.9%
66.7%9 pairs
90.6%

Chat suite

End-to-end question-and-answer cases. A pass means the answer was grounded in the right facts from the corpus. Failing case ids are published, not hidden.

32/35cases pass

Failing case ids: atlas_scope, strict_mode_hr, who_is_ana

Trends

The ten most recent releases from the v1 line on, oldest to newest, on an honest 0 to 100 percent axis. The complete history stays published in the repository. The dashed line is the continuous-integration gate that a release must clear to ship.

Measured at releaseBackfilled: transcribed from recorded runs rather than emitted by the harness at release time.CI gate
Extraction precision84.1%
Extraction recall94.6%
Verification agreement91.9%
Deduplication accuracy94.4%
Contradiction precision81.3%
Contradiction recall92.9%
Supersedes accuracy66.7%
Query-rewrite routing accuracy90.6%
Trend data for configuration mistral-default, Aggregate, all releases
ReleaseExtraction precisionExtraction recallVerification agreementDeduplication accuracyContradiction precisionContradiction recallSupersedes accuracyQuery-rewrite routing accuracy
v1.1.078.9%93.3%86.5%92.9%not measured100%not measurednot measured
v1.2.079.2%91.3%89.2%92.9%not measured100%not measurednot measured
v1.3.079.6%92%86.9%92.9%not measured100%not measurednot measured
v1.4.081.3%93.8%92.9%92.9%not measured100%not measurednot measured
v1.4.178.9%91.1%83.3%92.9%66.7%100%75%90.6%
v1.4.2-local78.9%91.1%83.3%92.9%66.7%100%75%90.6%
v1.5.082.7%93.9%92.6%94.4%75%100%75%90.6%
v1.6.082.3%90.4%91.7%94.4%82.4%100%66.7%84.4%
v1.7.184.2%95.8%88.9%94.4%81.3%92.9%66.7%90.6%
v1.8.084.1%94.6%91.9%94.4%81.3%92.9%66.7%90.6%

Notes from the releases

  • v1.4.2-localOpen the JSON file
    • Self-hosted configuration measurement (llama.cpp ff711 + bge-m3 on the operator's own hardware), thinking suppressed on every call. Advisory against the Mistral-measured gates: contradiction recall and query-rewrite accuracy sit below their floors; both zero-tolerance gates (injection, subject) pass. Not a release measurement.

Provenance

Each release, with the exact commit it was measured at, the harness version, the corpus sizes, and a direct link to its immutable JSON file. Published files are never edited after release. Read the data, not our summary of it.

  • Harness
    extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases

  • Harness
    extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases

  • Harness
    extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases

  • Harness
    extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 98 golden cases (48 english, 50 croatian) · 33 reconciliation pairs · 24 chat cases

  • v1.4.2-local2026-08-05Open the JSON filebe2254dc6e
    Harness
    extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration cogeto-offline

    Models: pipeline cogeto, answer cogeto, embedding bge-m3

    Corpus: 98 golden cases (48 english, 50 croatian) · 33 reconciliation pairs · 24 chat cases

    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 29 reconciliation pairs · 24 chat cases

  • Harness
    extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 29 reconciliation pairs · 24 chat cases

  • Harness
    extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 20 reconciliation pairs · 24 chat cases

  • Harness
    extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 20 reconciliation pairs · 24 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 76 golden cases (37 english, 39 croatian) · 20 reconciliation pairs · 24 chat cases

  • Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0006 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 76 golden cases (37 english, 39 croatian) · 20 reconciliation pairs · 27 chat cases

Back to cogeto.eu