Trust scores
The measured accuracy of Cogeto, per release
Cogeto publishes its own measured accuracy for every release, the same way a service publishes uptime, including the numbers that fall short of their targets. Here are the numbers, and here are the public data files behind them. Do not trust this chart: check the file.
Aggregate blends the per-language corpora. It is shown so a weak language can never hide inside an average: switch the selector to read each language on its own.
Every gate floor is set at the honest current value of the metric, never at a target the project has not reached, and floors only ratchet upward. Floors apply per language as well as in aggregate, so the gate you see here changes with the language you select.
Current scores
Extraction and reconciliation quality for the selected model configuration and language, measured against a hand-labeled golden corpus.
| Metric | Aggregate | CI gate |
|---|---|---|
| 84.1% | CI gate ≥ 77% CI gate: pass | |
| 94.6% | CI gate ≥ 91% CI gate: pass | |
| 91.9% | CI gate ≥ 86% CI gate: pass | |
| 94.4% | CI gate ≥ 92% CI gate: pass | |
| 81.3% | CI gate ≥ 54% CI gate: pass | |
| 92.9% | CI gate ≥ 83% CI gate: pass | |
| 66.7%9 pairs | CI gate ≥ 50% CI gate: pass | |
| 90.6% | CI gate ≥ 81% CI gate: pass |
Chat suite
End-to-end question-and-answer cases. A pass means the answer was grounded in the right facts from the corpus. Failing case ids are published, not hidden.
32/35cases pass
Failing case ids: atlas_scope, strict_mode_hr, who_is_ana
Trends
The ten most recent releases from the v1 line on, oldest to newest, on an honest 0 to 100 percent axis. The complete history stays published in the repository. The dashed line is the continuous-integration gate that a release must clear to ship.
| Release | Extraction precision | Extraction recall | Verification agreement | Deduplication accuracy | Contradiction precision | Contradiction recall | Supersedes accuracy | Query-rewrite routing accuracy |
|---|---|---|---|---|---|---|---|---|
| v1.1.0 | 78.9% | 93.3% | 86.5% | 92.9% | not measured | 100% | not measured | not measured |
| v1.2.0 | 79.2% | 91.3% | 89.2% | 92.9% | not measured | 100% | not measured | not measured |
| v1.3.0 | 79.6% | 92% | 86.9% | 92.9% | not measured | 100% | not measured | not measured |
| v1.4.0 | 81.3% | 93.8% | 92.9% | 92.9% | not measured | 100% | not measured | not measured |
| v1.4.1 | 78.9% | 91.1% | 83.3% | 92.9% | 66.7% | 100% | 75% | 90.6% |
| v1.4.2-local | 78.9% | 91.1% | 83.3% | 92.9% | 66.7% | 100% | 75% | 90.6% |
| v1.5.0 | 82.7% | 93.9% | 92.6% | 94.4% | 75% | 100% | 75% | 90.6% |
| v1.6.0 | 82.3% | 90.4% | 91.7% | 94.4% | 82.4% | 100% | 66.7% | 84.4% |
| v1.7.1 | 84.2% | 95.8% | 88.9% | 94.4% | 81.3% | 92.9% | 66.7% | 90.6% |
| v1.8.0 | 84.1% | 94.6% | 91.9% | 94.4% | 81.3% | 92.9% | 66.7% | 90.6% |
Notes from the releases
- v1.4.2-localOpen the JSON file
- Self-hosted configuration measurement (llama.cpp ff711 + bge-m3 on the operator's own hardware), thinking suppressed on every call. Advisory against the Mistral-measured gates: contradiction recall and query-rewrite accuracy sit below their floors; both zero-tolerance gates (injection, subject) pass. Not a release measurement.
Provenance
Each release, with the exact commit it was measured at, the harness version, the corpus sizes, and a direct link to its immutable JSON file. Published files are never edited after release. Read the data, not our summary of it.
- Harness
- extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases
- Harness
- extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases
- Harness
- extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases
- Harness
- extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 98 golden cases (48 english, 50 croatian) · 33 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
cogeto-offline Models: pipeline
cogeto, answercogeto, embeddingbge-m3Corpus: 98 golden cases (48 english, 50 croatian) · 33 reconciliation pairs · 24 chat cases
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 english, 44 croatian) · 29 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 english, 44 croatian) · 29 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 english, 44 croatian) · 20 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 english, 44 croatian) · 20 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 76 golden cases (37 english, 39 croatian) · 20 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0006 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 76 golden cases (37 english, 39 croatian) · 20 reconciliation pairs · 27 chat cases