Association, not independent accuracy: both axes contain the same Manifest transcript. Shared errors, clip difficulty, language and different reference lengths can all affect this relationship.
Hindi manifest · WER survival
How many clips retain agreement below your chosen WER threshold? Original Hindi manifest transcripts compared with Qwen-0.6B and Voxtral-24B.
Joining the same 500 clips…
Original manifest, not the ensemble
This page uses the stored corpus transcript as the hypothesis for both comparisons. You identified its producer as a Hindi-capable Parakeet; the saved metadata does not identify the exact checkpoint/version. These are machine transcripts, not human ground truth. The experimental romanized Hindi fine-tune is not used here.
Hindi only · 500 clips. Hindi was selected by SpeechBrain; Qwen reports English for 36 of its clips. No clips were removed after seeing ASR results. Existing ensemble pages remain unchanged.
Jump to WER survival chart ↓Linear association · raw WER
Rank association · ties averaged
Subtract each language’s mean first
What share of samples have WER below X%?
Hindi sample percentages, not average WER. Click a bar to browse matching clips.
Uses the coverage and quality filters above; uses the full Hindi cohort regardless of chart selection. Denominator: clips with both WER scores defined in each language. Undefined pairs are excluded and counted. Exact-threshold scores do not qualify; at X = 0, even exact matches are excluded.
Two disagreements, one dot per clip
Hover for scores. Click to listen. Drag a box to select clips; selection filters the browser below, not the headline statistics.
Axes show WER percentages, not accuracy. Values above 100% are retained. Changing axis scale never changes correlations. Use the table for keyboard-accessible clip selection.
Is the pattern consistent by language?
Click a row to filter. These are clip-level correlations, not correlations of language averages.
Can one disagreement screen for the other?
Choose a common WER threshold. “High” means WER ≥ threshold.
Click a quadrant to browse its clips. “Precision” and “recall” here only refer to the high-Voxtral-disagreement condition, not transcription correctness.
Average Voxtral disagreement by Qwen WER band
Click a band to inspect it. Fixed X bands; Y is mean per-clip WER, not pooled WER. Empty bands remain visible.
Follow the dots back to audio
| Language / clip | Duration | Manifest–Qwen WER | Manifest–Voxtral WER | Quality flags | Listen |
|---|
Methods, interpretation & limitations
X = normalized edit distance from Qwen-0.6B to Manifest, divided by Qwen words. Y = normalized edit distance from Voxtral-24B to the same Manifest output, divided by Voxtral words. We recover exact original manifest text, join clip IDs and audio SHA256, and preserve the existing Qwen/Voxtral outputs. Hindi joins also verify source file, episode, segment and start/end times. No new inference is performed.
Each clip has equal weight. Pearson uses raw WER; Spearman uses average ranks for ties. Within-language Pearson correlates residuals after subtracting each language’s X and Y means within the current selection. This removes group mean differences, not every language or clip confounder. Undefined WER pairs are excluded explicitly; constant or fewer than three paired scores give an undefined correlation. Scores are never capped or winsorized.
Hindi is within Voxtral’s advertised coverage. The coverage and quality filters are sensitivity checks, not a guarantee of correctness. Raw flagged outputs remain included by default. Clips may share an episode or speaker; no independence-based p-values or confidence intervals are claimed. The shared Manifest output makes this a descriptive screening analysis, not proof that Qwen predicts true ASR errors.
Metric definitions: SciPy Pearson · SciPy Spearman. Browser calculations are checked against SciPy for every language and preset scope.