Genauigkeitswerte
Die gemessene Genauigkeit von Cogeto, je Release
Cogeto veröffentlicht für jedes Release die eigene gemessene Genauigkeit, so wie ein Dienst seine Verfügbarkeit veröffentlicht, einschließlich der Zahlen, die ihre Zielwerte verfehlen. Hier sind die Zahlen, und hier sind die öffentlichen Datendateien dahinter. Vertrauen Sie nicht diesem Diagramm: Prüfen Sie die Datei.
Gesamt mischt die Korpora der einzelnen Sprachen. Es wird gezeigt, damit sich eine schwache Sprache nie in einem Durchschnitt verstecken kann: Stellen Sie den Wahlschalter um, um jede Sprache einzeln zu lesen.
Jede Gate-Untergrenze liegt beim ehrlichen aktuellen Wert der Metrik, nie bei einem Ziel, das das Projekt noch nicht erreicht hat, und Untergrenzen werden nur angehoben, nie gesenkt. Sie gelten je Sprache wie auch im Gesamtwert, das hier gezeigte Gate ändert sich deshalb mit der gewählten Sprache.
Aktuelle Werte
Extraktions- und Abgleichqualität für die gewählte Modellkonfiguration und Sprache, gemessen an einem von Hand annotierten Goldstandard-Korpus.
| Metrik | Gesamt | CI-Gate |
|---|---|---|
| 84.1% | CI-Gate ≥ 77% CI-Gate: pass | |
| 94.6% | CI-Gate ≥ 91% CI-Gate: pass | |
| 91.9% | CI-Gate ≥ 86% CI-Gate: pass | |
| 94.4% | CI-Gate ≥ 92% CI-Gate: pass | |
| 81.3% | CI-Gate ≥ 54% CI-Gate: pass | |
| 92.9% | CI-Gate ≥ 83% CI-Gate: pass | |
| 66.7%9 Paare | CI-Gate ≥ 50% CI-Gate: pass | |
| 90.6% | CI-Gate ≥ 81% CI-Gate: pass |
Chat-Testsuite
Frage-Antwort-Fälle von Ende zu Ende. Bestanden heißt: Die Antwort war in den richtigen Fakten aus dem Korpus verankert. Die IDs durchgefallener Fälle werden veröffentlicht, nicht versteckt.
32/35Fälle bestehen
IDs fehlgeschlagener Fälle: atlas_scope, strict_mode_hr, who_is_ana
Verläufe
Die zehn neuesten Releases ab der v1-Linie, vom ältesten zum neuesten, auf einer ehrlichen Achse von 0 bis 100 Prozent. Die vollständige Historie bleibt im Repository veröffentlicht. Die gestrichelte Linie ist das CI-Gate, das ein Release zum Ausliefern bestehen muss.
| Release | Extraktions-Precision | Extraktions-Recall | Verifikationsübereinstimmung | Deduplizierungsgenauigkeit | Widerspruchs-Precision | Widerspruchs-Recall | Ablösungsgenauigkeit | Genauigkeit des Query-Rewrite-Routings |
|---|---|---|---|---|---|---|---|---|
| v1.1.0 | 78.9% | 93.3% | 86.5% | 92.9% | nicht gemessen | 100% | nicht gemessen | nicht gemessen |
| v1.2.0 | 79.2% | 91.3% | 89.2% | 92.9% | nicht gemessen | 100% | nicht gemessen | nicht gemessen |
| v1.3.0 | 79.6% | 92% | 86.9% | 92.9% | nicht gemessen | 100% | nicht gemessen | nicht gemessen |
| v1.4.0 | 81.3% | 93.8% | 92.9% | 92.9% | nicht gemessen | 100% | nicht gemessen | nicht gemessen |
| v1.4.1 | 78.9% | 91.1% | 83.3% | 92.9% | 66.7% | 100% | 75% | 90.6% |
| v1.4.2-local | 78.9% | 91.1% | 83.3% | 92.9% | 66.7% | 100% | 75% | 90.6% |
| v1.5.0 | 82.7% | 93.9% | 92.6% | 94.4% | 75% | 100% | 75% | 90.6% |
| v1.6.0 | 82.3% | 90.4% | 91.7% | 94.4% | 82.4% | 100% | 66.7% | 84.4% |
| v1.7.1 | 84.2% | 95.8% | 88.9% | 94.4% | 81.3% | 92.9% | 66.7% | 90.6% |
| v1.8.0 | 84.1% | 94.6% | 91.9% | 94.4% | 81.3% | 92.9% | 66.7% | 90.6% |
Anmerkungen aus den Releases
- v1.4.2-localJSON-Datei öffnen
- Self-hosted configuration measurement (llama.cpp ff711 + bge-m3 on the operator's own hardware), thinking suppressed on every call. Advisory against the Mistral-measured gates: contradiction recall and query-rewrite accuracy sit below their floors; both zero-tolerance gates (injection, subject) pass. Not a release measurement.
Provenienz
Jedes Release mit dem exakten Commit, an dem gemessen wurde, der Version des Evaluations-Harness, den Korpusgrößen und einem direkten Link zu seiner unveränderlichen JSON-Datei. Veröffentlichte Dateien werden nach dem Release nie bearbeitet. Lesen Sie die Daten, nicht unsere Zusammenfassung.
- Harness
- extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 102 golden cases (50 englisch, 52 kroatisch) · 54 reconciliation pairs · 35 chat cases
- Harness
- extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 102 golden cases (50 englisch, 52 kroatisch) · 54 reconciliation pairs · 35 chat cases
- Harness
- extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 102 golden cases (50 englisch, 52 kroatisch) · 54 reconciliation pairs · 35 chat cases
- Harness
- extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 98 golden cases (48 englisch, 50 kroatisch) · 33 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
cogeto-offline Models: pipeline
cogeto, answercogeto, embeddingbge-m3Corpus: 98 golden cases (48 englisch, 50 kroatisch) · 33 reconciliation pairs · 24 chat cases
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 englisch, 44 kroatisch) · 29 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 englisch, 44 kroatisch) · 29 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 englisch, 44 kroatisch) · 20 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 86 golden cases (42 englisch, 44 kroatisch) · 20 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 76 golden cases (37 englisch, 39 kroatisch) · 20 reconciliation pairs · 24 chat cases
- Harness
- extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0006 · grader eval-coverage/v0001
- Configuration
mistral-default Models: pipeline
mistral-small-latest, answermistral-medium-latest, embeddingmistral-embedCorpus: 76 golden cases (37 englisch, 39 kroatisch) · 20 reconciliation pairs · 27 chat cases