Association, not independent accuracy: both axes contain the same Parakeet transcript. Shared errors, clip difficulty, language and different reference lengths can all affect this relationship.
Does disagreement travel?
When Parakeet disagrees with Qwen-0.6B, does it also disagree with Voxtral-24B?
Joining the same 9,000 clips…
Linear association · raw WER
Rank association · ties averaged
Subtract each language’s mean first
Two disagreements, one dot per clip
Hover for scores. Click to listen. Drag a box to select clips; selection filters the browser below, not the headline statistics.
Axes show WER percentages, not accuracy. Values above 100% are retained. Changing axis scale never changes correlations. Use the table for keyboard-accessible clip selection.
Is the pattern consistent by language?
Click a row to filter. These are clip-level correlations, not correlations of language averages.
Can one disagreement screen for the other?
Choose a common WER threshold. “High” means WER ≥ threshold.
Click a quadrant to browse its clips. “Precision” and “recall” here only refer to the high-Voxtral-disagreement condition, not transcription correctness.
Average Voxtral disagreement by Qwen WER band
Click a band to inspect it. Fixed X bands; Y is mean per-clip WER, not pooled WER. Empty bands remain visible.
Follow the dots back to audio
| Language / clip | Duration | Parakeet–Qwen WER | Parakeet–Voxtral WER | Quality flags | Listen |
|---|
Methods, interpretation & limitations
X = normalized edit distance from Qwen-0.6B to Parakeet, divided by Qwen words. Y = normalized edit distance from Voxtral-24B to the same Parakeet output, divided by Voxtral words. We join by exact clip ID and verify the Parakeet and Qwen transcripts are identical across both source datasets. No new inference is performed.
Each clip has equal weight. Pearson uses raw WER; Spearman uses average ranks for ties. Within-language Pearson correlates residuals after subtracting each language’s X and Y means within the current selection. This removes group mean differences, not every language or clip confounder. Undefined WER pairs are excluded explicitly; constant or fewer than three paired scores give an undefined correlation. Scores are never capped or winsorized.
Twelve corpus languages are outside Voxtral’s advertised coverage. The coverage and quality filters are sensitivity checks, not a guarantee of correctness. Raw flagged outputs remain included by default. Clips may share an episode or speaker; no independence-based p-values or confidence intervals are claimed. The shared Parakeet output makes this a descriptive screening analysis, not proof that Qwen predicts true ASR errors.
Metric definitions: SciPy Pearson · SciPy Spearman. Browser calculations are checked against SciPy for every language and preset scope.