Model-based language verification
Hindi, under the microscope.
Listen to SpeechBrain-selected Hindi podcast clips and inspect the community Parakeet candidate alongside Qwen3-ASR-0.6B and Voxtral-Small-24B.
Experimental · script mismatch, not an accuracy ranking
The community Hinglish Parakeet candidate emits SLP1-like Romanized Hindi. Qwen and Voxtral usually write Hindi in Devanagari. The same spoken words can therefore appear completely different to a word-error calculation.
Raw cross-script WER is a diagnostic, not comparable recognition accuracy. These 500 clips are excluded from the existing 18-language leaderboard and correlation charts. Raw model output is preserved exactly; no speculative transliteration is applied. Neither Qwen nor Voxtral is human ground truth.
The diagnostic normalization lowercases text, which loses sound distinctions in this case-sensitive Romanization. The transcripts below retain the original case and control tags.
At most five clips per episode
Same clip for all three models
SLP1-like output; raw text retained
Find a clip
Search in Hindi, Latin text, or clip ID.
Disagreement ordering uses normalized raw text, without transliteration. This can primarily reflect script choice.
Choose a clip
Playback unavailable. Try downloading this clip.
Community Parakeet candidate
Raw · SLP1-likeQwen3-ASR-0.6B
Model-generated transcriptVoxtral-Small-24B
Pseudo-reference · not ground truthShow raw diagnostic WER for this clip
Raw text, no script conversion. WER = (substitutions + deletions + insertions) ÷ reference words. It can exceed 100%; this is not an accuracy percentage.
| Hypothesis → reference | Raw WER |
|---|
Clip provenance
Show corpus-level raw diagnostic WER
These values are intentionally not headline scores. Cross-script comparisons are not comparable to the native-script 18-language evaluation.
| Hypothesis → reference | Pooled raw WER |
|---|
Models & provenance
A community candidate, not a verified Hindi accuracy winner. Exact base-checkpoint lineage and author transliteration recipe are undocumented.
What was held constant
Selection used local podcast good-audio segments, DNSMOS ≥ 3.8, 3–30 second eligibility, and at most five candidates per episode. SpeechBrain's pinned VoxLingua107 model verified each accepted clip's top language as Hindi. No source-manifest or old evaluator text was supplied as an inference prompt.