Lower means closer agreement
Every language. Every word.
Character-level disagreement
Disagreement by language
One pairwise WER. Two transcripts.
Click a language to explore its clips. Qwen-0.6B supplies the word-count denominator.
What the score means
Agreement, not an accuracy ranking.
Pairwise WER counts the word edits between the two transcripts. We normalize both outputs and divide pooled edits by the number of Qwen-0.6B words.
This view compares Parakeet ensemble directly with Qwen3-ASR-0.6B. Neither is human ground truth. Lower disagreement does not tell us which model is correct.
Hear the difference
for yourself.
Listen to the same audio, compare both transcripts, and inspect every word-level edit.
Why batched beam-8? Prior local experiments found batched mALSD-8 offered nearly the accuracy of conventional beam-8 at much higher throughput. That earlier benchmark is separate from this corpus.
The language ledger
Exact rates and the checkpoint behind each language.
| Language | Clips | Pairwise WER ↓ | Pairwise CER ↓ | Exact match | Parakeet checkpoint |
|---|
A distribution, not just an average
Loading clip distribution…
Filtered clip counts by pairwise WER bucket. WER can exceed 100%; it is not an accuracy percentage.
Clip explorer
Choose a clip to start listening.
How to read this evaluation
Reproducibility starts with the reference.
Best per language
Qwen3-ASR-0.6B
The two systems transcribed the same frozen clips. Qwen-0.6B supplies the computational reference for pairwise WER; it is not assigned an accuracy score.
A model for every language
Checkpoint provenance for the Parakeet ensemble.