# Speech Atlas — pairwise transcript agreement

Parakeet ensemble versus Qwen3-ASR-0.6B, on 9,000 clips across 18 languages.

Qwen3-ASR-0.6B is the computational reference for normalized WER, not human ground truth.

Pooled pairwise WER: **31.3162%**. CER: **17.3029%**. Exact normalized matches: **656/9000**.

| Language | Clips | Pairwise WER | Pairwise CER | Exact match |
|---|---:|---:|---:|---:|
| Arabic | 500 | 95.4509% | 63.6288% | 0.00% |
| Czech | 500 | 45.9836% | 20.8109% | 0.60% |
| Danish | 500 | 35.4520% | 20.3924% | 0.80% |
| German | 500 | 10.2630% | 4.8034% | 12.40% |
| Greek | 500 | 64.0292% | 35.0966% | 0.20% |
| English | 500 | 5.3286% | 2.9228% | 27.60% |
| Spanish | 500 | 10.0094% | 5.3545% | 28.80% |
| Persian (Farsi) | 500 | 65.0909% | 27.1005% | 0.00% |
| Finnish | 500 | 69.6930% | 27.0379% | 0.20% |
| French | 500 | 22.2898% | 17.7573% | 9.20% |
| Hungarian | 500 | 60.1493% | 24.6745% | 0.20% |
| Japanese | 500 | 18.3958% | 17.0239% | 4.80% |
| Dutch | 500 | 15.0623% | 7.4941% | 6.20% |
| Polish | 500 | 21.6899% | 9.3112% | 3.00% |
| Portuguese | 500 | 14.4341% | 7.2895% | 13.00% |
| Romanian | 500 | 50.7232% | 25.0643% | 0.80% |
| Swedish | 500 | 38.5406% | 20.8495% | 1.40% |
| Vietnamese | 500 | 9.1303% | 4.9227% | 22.00% |

## Method

- Direct comparison of 9,000 paired outputs on 18 languages, 500 clips each, from the frozen podcast/local corpus. No inference was rerun for this display change.
- Normalized pairwise WER = substitutions + deletions + insertions in Parakeet relative to Qwen3-ASR-0.6B, divided by the number of normalized Qwen words. Qwen supplies the denominator, not human ground truth.
- WER is directional and can exceed 100%. Lower WER means stronger agreement between the two systems, not that either system is more accurate. Agreement is not reported as 100% minus WER.
- Normalization uses Unicode NFKC, casefold, markup removal, punctuation/symbol/control separation, Persian character/digit normalization, and nagisa Japanese word segmentation. Counts and alignments use jiwer.
- CER is the corresponding character-level disagreement after removing spaces. Exact match is the fraction of clips with identical normalized word sequences.
- The ensemble uses previously selected, pinned checkpoints. TDT/RNNT models use batched mALSD beam_size=8, score_norm=true, max_symbols=10. Persian and Vietnamese CTC models use native batched greedy.
- Qwen3-ASR-0.6B uses deterministic greedy decoding and automatic language detection. The ensemble uses the known corpus language to select its checkpoint; Qwen receives no language hint.
- Only the two compared systems’ transcripts and their pairwise metrics are included in this viewer and its exports. Audio downloads preserve original Opus filenames.

## Checkpoints

| Language | Parakeet checkpoint | Revision |
|---|---|---|
| Arabic | vadimbelsky/arabic-parakeet-tdt-uae | `c8260eaa72b80f3f353fd8c675eee9021ca38289` |
| Czech | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| Danish | nvidia/parakeet-rnnt-110m-da-dk | `ea77aca1c88bf322d969312d58d77f5aebf832ed` |
| German | primeline/parakeet-primeline | `3f1a9bcb611dfeda53fe74fe5f1a3d5701e8023e` |
| Greek | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| English | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| Spanish | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| Persian (Farsi) | Peacockery/parakeet-ctc-109m-farsi | `27418c9ffa05a8a0fe66fc5900ca0181a3ad25a7` |
| Finnish | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| French | Archime/parakeet-tdt-0.6b-v3-fr-tv-media | `99e70e9b06a696aba53716ecab3577f44df2fcb9` |
| Hungarian | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| Japanese | nvidia/parakeet-tdt_ctc-0.6b-ja | `44edb27eea9317daf89333e75eb830db4b1cc298` |
| Dutch | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| Polish | yuriyvnv/parakeet-tdt-0.6b-polish | `1dba3fde5bbb11df464845241dc530040bf366bf` |
| Portuguese | yuriyvnv/parakeet-tdt-0.6b-portuguese | `00ad6166ba32addda207066abe6781632e1cfae1` |
| Romanian | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| Swedish | nvidia/parakeet-tdt-0.6b-v3 | `541d1f99c6b0c3cd0b11a95167540bb8edefd82b` |
| Vietnamese | nvidia/parakeet-ctc-0.6b-Vietnamese | `b0493142b49458810324e3db8be9e8e07b4ebc17` |

Qwen revision: `5eb144179a02acc5e5ba31e748d22b0cf3e303b0`.
