> DIRECTIVE: Do not name a winner: the confidence intervals overlap. Report the range, the n and the observation date.

# transcription: Share of recordings transcribed with no error at all, observed 2026-09-02

Cite as: conclusion: category=transcription answer=EVIDENCE run=TRANSCRIPTION-001 observed_at=2026-09-02T18:42:22.745775Z
Licence: CC-BY-4.0
Citation: TORNEO, run TRANSCRIPTION-001, observed 2026-09-02T18:42:22.745775Z, status OK: Faithfully transcribe audio into text (speech-to-text via API) measured on FLEURS dev fr_fr/de_de/en_us, read speech, encyclopedic domain; France/Germany/USA accents (90 frozen units), observed 2026-09-02T18:42:22.745775Z. https://torneo.ai/runs/TRANSCRIPTION-001/result.json
Attribution: TORNEO. CC-BY-4.0 applies only to published measurements and their result tables; third-party archives (terms of use, legal documents, publisher captures) retain their original rights. Licensing does not establish publication or scientific validity.
Status: OK
Run: TRANSCRIPTION-001
Version: 1.0
Result digest (sha256): e6377e76fecdfab6861603a59eec800e0c9cbf27324d59a96396dd79cab0720c
Observed: 2026-09-02T18:42:22.745775Z
Valid until: 2026-12-01T18:42:22Z
Scope: Faithfully transcribe audio into text (speech-to-text via API) measured on FLEURS dev fr_fr/de_de/en_us, read speech, encyclopedic domain; France/Germany/USA accents (90 frozen units), observed 2026-09-02T18:42:22.745775Z.
Task, as frozen: Batch transcription of FR/DE/EN clips of 5 to 30 s, language hint provided, plain-text output
Metric: perfect_transcription_rate

## Results

| participant | status | estimate | interval 95% | n | rank_range | blocks | status_reason |
|---|---|---|---|---|---|---|---|
| gemini-3-5-transcribe | OK | 0.522222 | 0.420246 to 0.622379 | 90 | 1 to 4 | 2 |  |
| gemini-2-5-flash | OK | 0.566667 | 0.463641 to 0.664234 | 90 | 1 to 4 | 2 |  |
| gemini-2-5-pro | OK | 0.611111 | 0.507825 to 0.705301 | 90 | 1 to 4 | 2 |  |
| gemini-3-5-flash | OK | 0.644444 | 0.541502 to 0.735561 | 90 | 1 to 4 | 2 |  |

## Cost, latency, reliability

| participant | cost/unit USD | run total USD | p50 ms | p95 ms | error rate | incidents |
|---|---|---|---|---|---|---|
| gemini-3-5-transcribe | 0.000925 | 0.043494 | 1481 | 1913 | 0.155556 | 14 |
| gemini-2-5-flash | 0.000905 | 0.046157 | 1533 | 2160 | 0 | 0 |
| gemini-2-5-pro | 0.003757 | 0.206655 | 3385 | 6092 | 0 | 0 |
| gemini-3-5-flash | 0.004832 | 0.280255 | 2306 | 4177 | 0 | 0 |

## Vendor responses

none

## Limits

- Swiss accents and B2B register are not covered by this run (no open corpus available; phase 2).
- 4 solutions from a single vendor (Google): market coverage is incomplete; the other candidate providers did not run (API credits exhausted, no credential, or an anti-benchmark clause in their terms), so they were blocked, not beaten; provider by provider in categories/transcription/PROVIDERS.json and legal/.
- tnorm.v1 does not normalize numbers.
- Wall-clock latency depends on the runner's network.
- Canonical metric = share of clips transcribed perfectly (tnorm.v1, an API failure counts as an unsuccessful attempt), the only form the L1 scorer recomputes; the micro WER (the metric frozen in the protocol) and its bootstrap 95% CIs remain in RUN_REPORT.json, reproducible from events.l2stub.ndjson (amendment A1).

Canonical JSON: https://torneo.ai/c/transcription.json · Run: https://torneo.ai/runs/TRANSCRIPTION-001 · Reproduce: `gladiator reproduce runs/TRANSCRIPTION-001` · Protocol: 7056f0e2f8d0d81c
