Leaderboard ng Mongolian speech-to-text
Ihambing ang tasa ng error sa salita, bilis, at presyo ng mga provider ng speech-to-text para sa Mongolian. Bawat numero ay sinusukat sa tunay na Mongolian audio — hindi marketing.
Datasets used
Public speech corpora used for every provider run. Open a source URL for the upstream dataset card, license, and download.
- Common Voice 24 (MN)›https://huggingface.co/datasets/btsee/common-voices-24-mn
- Shunya Labs Mongolian Speech›https://huggingface.co/datasets/shunyalabs/mongolian-speech-dataset
- Common Voice 20 (MN)›https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized
- Modern Voice›Recorded first-party for this benchmark — audio not distributed, metrics only.
WER vs bilis
Bawat tuldok ay isang provider. Ang ideal na posisyon ay sa kaliwang baba (kaunting error, mabilis). Sinusukat sa tunay na Mongolian audio.
Paghahambing ng provider
Ayusin ang anumang column. I-click ang Detalye para sa paghihimay per sample, presyo at metodolohiya ng isang provider.
| Provider | WER ↑ | CER | Katumpakan | Bilis | Latency | Presyo / 1000 min | |
|---|---|---|---|---|---|---|---|
Speechmatics | 14.4% | 7.4% | 85.6% | 0.51× | 13.8s | $8.5 | Detalye |
Duudlaga FlowAmin | 14.6% | 7.6% | 85.4% | 1.47× | 5.2s | $19.4 | Detalye |
Chimege | 14.6% | 6.6% | 85.4% | 1.51× | 4.7s | $11.1 | Detalye |
Google STT | 16.9% | 7.4% | 83.1% | 1.98× | 3.7s | $16 | Detalye |
Whisper large-v2 MN | 20.7% | 8.3% | 79.3% | 0.53× | 13.5s | $0 | Detalye |
Azure Speech | 20.8% | 9.7% | 79.2% | 0.39× | 17.3s | $16.7 | Detalye |
Whisper medium MNOverlap | 22.5% | 8.9% | 77.5% | 2.7× | 2.7s | $0 | Detalye |
SeamlessM4T v2 | 24% | 11.1% | 76% | 0.39× | 26.3s | $0 | Detalye |
Whisper turbo MNOverlap | 25.7% | 9.2% | 74.3% | 2.43× | 3.0s | $0 | Detalye |
ElevenLabs Scribe v1 | 27.1% | 10.8% | 72.9% | 2.14× | 3.3s | $6.7 | Detalye |
ElevenLabs Scribe v2 | 27.1% | 10.6% | 72.9% | 1.51× | 4.7s | $3.7 | Detalye |
Gemma 4 | 30.2% | 14.7% | 69.8% | 1× | 7.3s | $0 | Detalye |
OmniASR LLM-1B | 33.8% | 14.2% | 66.2% | 0.27× | 26.6s | $0 | Detalye |
Whisper large-v3 MNOverlap | 39.4% | 14.9% | 60.6% | 1.57× | 4.5s | $0 | Detalye |
MMS-1B-all | 41.7% | 12.3% | 58.3% | 19.2× | 0.4s | $0 | Detalye |
wav2vec2 XLSR-53 MNOverlap | 45% | 16.4% | 55% | 42.32× | 0.2s | $0 | Detalye |
Dolphin small | 49.1% | 19.2% | 50.9% | 1.73× | 4.1s | $0 | Detalye |
OmniASR CTC-1B | 51.2% | 15.5% | 48.8% | 0.6× | 11.7s | $0 | Detalye |
GPT-4o Transcribe | 53% | 27.6% | 47% | 2.11× | 3.7s | $6 | Detalye |
Moonshine MNOverlap | 58.8% | 45.2% | 41.2% | 8.44× | 0.8s | $0 | Detalye |
Whisper large-v3 | 89.5% | 37.4% | 10.5% | 1.29× | 6.0s | $0 | Detalye |
Gladia | 90.1% | 38.9% | 9.9% | 0.75× | 9.4s | $10.2 | Detalye |
Whisper large-v3-turbo | 99% | 54.4% | 1% | 0.91× | 8.7s | $0 | Detalye |
Qwen3-ASR-Flash | 103.7% | 86.4% | -3.7% | 2.19× | 3.2s | $2.1 | Detalye |
OpenAI Whisper | 105.2% | 63.7% | -5.2% | 1.5× | 4.5s | $6 | Detalye |
VibeVoice ASR | 107% | 61% | -7% | 0.53× | 12.1s | $0 | Detalye |
Gemini Flash | Paparating na ang benchmark | ||||||
Lahat ng resulta ay sinusukat sa tunay na Mongolian dataset. Ang bilis ay × real-time (mas mataas = mas mabilis). Mas mababang WER ay mas mainam.
Overlap— Models with this badge were trained on data that overlaps the benchmark corpora (Mongolian Common Voice — 113 of 173 samples — or the public Shunya Labs corpus). Scores on overlapping corpora are inflated by memorization; judge these models by the per-dataset table on their details page, especially the corpora they were NOT trained on.
Duudlaga Voice Set
92 samples · 10.2 min · 28 systemsA second corpus recorded first-party for this benchmark, covering what the public Mongolian datasets barely contain: modern loanwords, English/Mongolian code-switching, numbers, dates, and commands. Every system is measured on the same 92 recordings.
| Provider | WER | CER | Katumpakan | Bilis |
|---|---|---|---|---|
1Duudlaga FlowAmin | 9.5% | 5.3% | 90.5% | 1.27× |
2Google STT | 11.7% | 5.5% | 88.3% | 2.28× |
3Chimege | 14.5% | 6.8% | 85.5% | 1.48× |
4Speechmatics | 15.3% | 7.6% | 84.7% | 0.5× |
5Gemini 3.5 Flash | 19.5% | 11.4% | 80.5% | 0.06× |
6SeamlessM4T v2Open source | 22.9% | 10.4% | 77.1% | 4.32× |
7Whisper large-v2 MNOpen source | 23.2% | 10.9% | 76.8% | 0.55× |
8ElevenLabs Scribe v1 | 23.4% | 9.3% | 76.6% | 2.09× |
9ElevenLabs Scribe v2 | 24.6% | 9.8% | 75.4% | 1.47× |
10Azure Speech | 26.1% | 10.9% | 73.9% | 0.32× |
11Gemma 4Open source | 27.7% | 13.2% | 72.3% | 1.04× |
12OmniASR LLM-1BOpen source | 34.5% | 14.5% | 65.5% | 0.27× |
13Whisper medium MNOpen sourceOverlap | 35.1% | 14.8% | 64.9% | 2.71× |
14Whisper turbo MNOpen sourceOverlap | 37.6% | 14.1% | 62.4% | 2.59× |
15Whisper large-v3 MNOpen sourceOverlap | 41% | 17.1% | 59% | 1.58× |
16GPT-4o Transcribe | 44.3% | 23.7% | 55.7% | 2.62× |
17MMS-1B-allOpen source | 45.3% | 15.1% | 54.7% | 20.52× |
18Dolphin smallOpen source | 46.1% | 19% | 53.9% | 1.69× |
19W2v-BERT 2.0 MNOpen sourceOverlap | 46.3% | 18.2% | 53.7% | 28.12× |
20OmniASR CTC-1BOpen source | 54.9% | 18.9% | 45.1% | 0.59× |
21wav2vec2 XLSR-53 MNOpen sourceOverlap | 56.3% | 23.9% | 43.7% | 45.7× |
22Whisper large-v3Open source | 87.3% | 36.6% | 12.7% | 1.55× |
23Gladia | 88.1% | 39.9% | 11.9% | 0.74× |
24Moonshine MNOpen sourceOverlap | 96.2% | 75.7% | 3.8% | 8.36× |
25Whisper large-v3-turboOpen source | 98.1% | 57.1% | 1.9% | 1.19× |
26Qwen3-ASR-Flash | 103.6% | 89.9% | -3.6% | 2.07× |
27OpenAI Whisper | 104.6% | 63.8% | -4.6% | 1.32× |
28VibeVoice ASROpen source | 110.1% | 64% | -10.1% | 0.4× |
Recorded first-party for this benchmark — audio not distributed, metrics only.