Experimental · script mismatch, not an accuracy ranking

The community Hinglish Parakeet candidate emits SLP1-like Romanized Hindi. Qwen and Voxtral usually write Hindi in Devanagari. The same spoken words can therefore appear completely different to a word-error calculation.

Raw cross-script WER is a diagnostic, not comparable recognition accuracy. These 500 clips are excluded from the existing 18-language leaderboard and correlation charts. Raw model output is preserved exactly; no speculative transliteration is applied. Neither Qwen nor Voxtral is human ground truth.

The diagnostic normalization lowercases text, which loses sound distinctions in this case-sensitive Romanization. The transcripts below retain the original case and control tags.

Loading the Hindi listening room…