Leaderboard ng Mongolian speech-to-text

Ihambing ang tasa ng error sa salita, bilis, at presyo ng mga provider ng speech-to-text para sa Mongolian. Bawat numero ay sinusukat sa tunay na Mongolian audio — hindi marketing.

Datasets used

Public speech corpora used for every provider run. Open a source URL for the upstream dataset card, license, and download.

WER vs bilis

Bawat tuldok ay isang provider. Ang ideal na posisyon ay sa kaliwang baba (kaunting error, mabilis). Sinusukat sa tunay na Mongolian audio.

Google STTDuudlaga FlowSpeechmaticsChimegeOpenAI WhisperGPT-4o TranscribeAzure SpeechGladiaGemma 4Whisper large-v3Whisper large-v3-turboWhisper large-v2 MNSeamlessM4T v2MMS-1B-allMoonshine MNwav2vec2 XLSR-53 MNWhisper turbo MNWhisper medium MNElevenLabs Scribe v1ElevenLabs Scribe v2Qwen3-ASR-FlashDolphin smallOmniASR CTC-1BOmniASR LLM-1BVibeVoice ASRWhisper large-v3 MN

Paghahambing ng provider

Ayusin ang anumang column. I-click ang Detalye para sa paghihimay per sample, presyo at metodolohiya ng isang provider.

ProviderWERCERKatumpakanBilisLatencyPresyo / 1000 min
Speechmatics
14.4%7.4%85.6%0.51×13.8s$8.5Detalye
Duudlaga FlowAmin
14.6%7.6%85.4%1.47×5.2s$19.4Detalye
Chimege
14.6%6.6%85.4%1.51×4.7s$11.1Detalye
Google STT
16.9%7.4%83.1%1.98×3.7s$16Detalye
Whisper large-v2 MN
20.7%8.3%79.3%0.53×13.5s$0Detalye
Azure Speech
20.8%9.7%79.2%0.39×17.3s$16.7Detalye
Whisper medium MNOverlap
22.5%8.9%77.5%2.7×2.7s$0Detalye
SeamlessM4T v2
24%11.1%76%0.39×26.3s$0Detalye
Whisper turbo MNOverlap
25.7%9.2%74.3%2.43×3.0s$0Detalye
ElevenLabs Scribe v1
27.1%10.8%72.9%2.14×3.3s$6.7Detalye
ElevenLabs Scribe v2
27.1%10.6%72.9%1.51×4.7s$3.7Detalye
Gemma 4
30.2%14.7%69.8%1×7.3s$0Detalye
OmniASR LLM-1B
33.8%14.2%66.2%0.27×26.6s$0Detalye
Whisper large-v3 MNOverlap
39.4%14.9%60.6%1.57×4.5s$0Detalye
MMS-1B-all
41.7%12.3%58.3%19.2×0.4s$0Detalye
wav2vec2 XLSR-53 MNOverlap
45%16.4%55%42.32×0.2s$0Detalye
Dolphin small
49.1%19.2%50.9%1.73×4.1s$0Detalye
OmniASR CTC-1B
51.2%15.5%48.8%0.6×11.7s$0Detalye
GPT-4o Transcribe
53%27.6%47%2.11×3.7s$6Detalye
Moonshine MNOverlap
58.8%45.2%41.2%8.44×0.8s$0Detalye
Whisper large-v3
89.5%37.4%10.5%1.29×6.0s$0Detalye
Gladia
90.1%38.9%9.9%0.75×9.4s$10.2Detalye
Whisper large-v3-turbo
99%54.4%1%0.91×8.7s$0Detalye
Qwen3-ASR-Flash
103.7%86.4%-3.7%2.19×3.2s$2.1Detalye
OpenAI Whisper
105.2%63.7%-5.2%1.5×4.5s$6Detalye
VibeVoice ASR
107%61%-7%0.53×12.1s$0Detalye
Gemini Flash
Paparating na ang benchmark

Lahat ng resulta ay sinusukat sa tunay na Mongolian dataset. Ang bilis ay × real-time (mas mataas = mas mabilis). Mas mababang WER ay mas mainam.

OverlapModels with this badge were trained on data that overlaps the benchmark corpora (Mongolian Common Voice — 113 of 173 samples — or the public Shunya Labs corpus). Scores on overlapping corpora are inflated by memorization; judge these models by the per-dataset table on their details page, especially the corpora they were NOT trained on.

Duudlaga Voice Set

92 samples · 10.2 min · 28 systems

A second corpus recorded first-party for this benchmark, covering what the public Mongolian datasets barely contain: modern loanwords, English/Mongolian code-switching, numbers, dates, and commands. Every system is measured on the same 92 recordings.

ProviderWERCERKatumpakanBilis
1Duudlaga FlowAmin
9.5%5.3%90.5%1.27×
2Google STT
11.7%5.5%88.3%2.28×
3Chimege
14.5%6.8%85.5%1.48×
4Speechmatics
15.3%7.6%84.7%0.5×
5Gemini 3.5 Flash
19.5%11.4%80.5%0.06×
6SeamlessM4T v2Open source
22.9%10.4%77.1%4.32×
7Whisper large-v2 MNOpen source
23.2%10.9%76.8%0.55×
8ElevenLabs Scribe v1
23.4%9.3%76.6%2.09×
9ElevenLabs Scribe v2
24.6%9.8%75.4%1.47×
10Azure Speech
26.1%10.9%73.9%0.32×
11Gemma 4Open source
27.7%13.2%72.3%1.04×
12OmniASR LLM-1BOpen source
34.5%14.5%65.5%0.27×
13Whisper medium MNOpen sourceOverlap
35.1%14.8%64.9%2.71×
14Whisper turbo MNOpen sourceOverlap
37.6%14.1%62.4%2.59×
15Whisper large-v3 MNOpen sourceOverlap
41%17.1%59%1.58×
16GPT-4o Transcribe
44.3%23.7%55.7%2.62×
17MMS-1B-allOpen source
45.3%15.1%54.7%20.52×
18Dolphin smallOpen source
46.1%19%53.9%1.69×
19W2v-BERT 2.0 MNOpen sourceOverlap
46.3%18.2%53.7%28.12×
20OmniASR CTC-1BOpen source
54.9%18.9%45.1%0.59×
21wav2vec2 XLSR-53 MNOpen sourceOverlap
56.3%23.9%43.7%45.7×
22Whisper large-v3Open source
87.3%36.6%12.7%1.55×
23Gladia
88.1%39.9%11.9%0.74×
24Moonshine MNOpen sourceOverlap
96.2%75.7%3.8%8.36×
25Whisper large-v3-turboOpen source
98.1%57.1%1.9%1.19×
26Qwen3-ASR-Flash
103.6%89.9%-3.6%2.07×
27OpenAI Whisper
104.6%63.8%-4.6%1.32×
28VibeVoice ASROpen source
110.1%64%-10.1%0.4×

Recorded first-party for this benchmark — audio not distributed, metrics only.