Mongolian Speech-to-Text Leaderboard

Compare word error rate, speed, and pricing across speech-to-text providers for Mongolian. Every number is measured on real Mongolian audio — not marketing.

Datasets used

Public speech corpora used for every provider run. Open a source URL for the upstream dataset card, license, and download.

WER vs speed

Each point is one provider. The sweet spot is the bottom-right (low error, fast). Measured on real Mongolian audio.

Google STTDuudlaga FlowSpeechmaticsChimegeOpenAI WhisperGPT-4o TranscribeAzure SpeechGladiaGemma 4Whisper large-v3Whisper large-v3-turboWhisper large-v2 MNSeamlessM4T v2MMS-1B-allMoonshine MNwav2vec2 XLSR-53 MNWhisper turbo MNWhisper medium MNElevenLabs Scribe v1ElevenLabs Scribe v2Qwen3-ASR-FlashDolphin smallOmniASR CTC-1BOmniASR LLM-1BVibeVoice ASRWhisper large-v3 MN

Provider comparison

Sort any column. Click Details for the full per-sample breakdown, pricing, and methodology for a provider.

ProviderWERCERAccuracySpeedLatencyPrice / 1k min
Speechmatics
14.4%7.4%85.6%0.51×13.8s$8.5Details
Duudlaga FlowOurs
14.6%7.6%85.4%1.47×5.2s$19.4Details
Chimege
14.6%6.6%85.4%1.51×4.7s$11.1Details
Google STT
16.9%7.4%83.1%1.98×3.7s$16Details
Whisper large-v2 MN
20.7%8.3%79.3%0.53×13.5s$0Details
Azure Speech
20.8%9.7%79.2%0.39×17.3s$16.7Details
Whisper medium MNOverlap
22.5%8.9%77.5%2.7×2.7s$0Details
SeamlessM4T v2
24%11.1%76%0.39×26.3s$0Details
Whisper turbo MNOverlap
25.7%9.2%74.3%2.43×3.0s$0Details
ElevenLabs Scribe v1
27.1%10.8%72.9%2.14×3.3s$6.7Details
ElevenLabs Scribe v2
27.1%10.6%72.9%1.51×4.7s$3.7Details
Gemma 4
30.2%14.7%69.8%1×7.3s$0Details
OmniASR LLM-1B
33.8%14.2%66.2%0.27×26.6s$0Details
Whisper large-v3 MNOverlap
39.4%14.9%60.6%1.57×4.5s$0Details
MMS-1B-all
41.7%12.3%58.3%19.2×0.4s$0Details
wav2vec2 XLSR-53 MNOverlap
45%16.4%55%42.32×0.2s$0Details
Dolphin small
49.1%19.2%50.9%1.73×4.1s$0Details
OmniASR CTC-1B
51.2%15.5%48.8%0.6×11.7s$0Details
GPT-4o Transcribe
53%27.6%47%2.11×3.7s$6Details
Moonshine MNOverlap
58.8%45.2%41.2%8.44×0.8s$0Details
Whisper large-v3
89.5%37.4%10.5%1.29×6.0s$0Details
Gladia
90.1%38.9%9.9%0.75×9.4s$10.2Details
Whisper large-v3-turbo
99%54.4%1%0.91×8.7s$0Details
Qwen3-ASR-Flash
103.7%86.4%-3.7%2.19×3.2s$2.1Details
OpenAI Whisper
105.2%63.7%-5.2%1.5×4.5s$6Details
VibeVoice ASR
107%61%-7%0.53×12.1s$0Details
Gemini Flash
Benchmark coming soon

All results are measured on real Mongolian datasets. Speed is × realtime (higher is faster). Lower WER is better.

OverlapModels with this badge were trained on data that overlaps the benchmark corpora (Mongolian Common Voice — 113 of 173 samples — or the public Shunya Labs corpus). Scores on overlapping corpora are inflated by memorization; judge these models by the per-dataset table on their details page, especially the corpora they were NOT trained on.

Duudlaga Voice Set

92 samples · 10.2 min · 28 systems

A second corpus recorded first-party for this benchmark, covering what the public Mongolian datasets barely contain: modern loanwords, English/Mongolian code-switching, numbers, dates, and commands. Every system is measured on the same 92 recordings.

ProviderWERCERAccuracySpeed
1Duudlaga FlowOurs
9.5%5.3%90.5%1.27×
2Google STT
11.7%5.5%88.3%2.28×
3Chimege
14.5%6.8%85.5%1.48×
4Speechmatics
15.3%7.6%84.7%0.5×
5Gemini 3.5 Flash
19.5%11.4%80.5%0.06×
6SeamlessM4T v2Open source
22.9%10.4%77.1%4.32×
7Whisper large-v2 MNOpen source
23.2%10.9%76.8%0.55×
8ElevenLabs Scribe v1
23.4%9.3%76.6%2.09×
9ElevenLabs Scribe v2
24.6%9.8%75.4%1.47×
10Azure Speech
26.1%10.9%73.9%0.32×
11Gemma 4Open source
27.7%13.2%72.3%1.04×
12OmniASR LLM-1BOpen source
34.5%14.5%65.5%0.27×
13Whisper medium MNOpen sourceOverlap
35.1%14.8%64.9%2.71×
14Whisper turbo MNOpen sourceOverlap
37.6%14.1%62.4%2.59×
15Whisper large-v3 MNOpen sourceOverlap
41%17.1%59%1.58×
16GPT-4o Transcribe
44.3%23.7%55.7%2.62×
17MMS-1B-allOpen source
45.3%15.1%54.7%20.52×
18Dolphin smallOpen source
46.1%19%53.9%1.69×
19W2v-BERT 2.0 MNOpen sourceOverlap
46.3%18.2%53.7%28.12×
20OmniASR CTC-1BOpen source
54.9%18.9%45.1%0.59×
21wav2vec2 XLSR-53 MNOpen sourceOverlap
56.3%23.9%43.7%45.7×
22Whisper large-v3Open source
87.3%36.6%12.7%1.55×
23Gladia
88.1%39.9%11.9%0.74×
24Moonshine MNOpen sourceOverlap
96.2%75.7%3.8%8.36×
25Whisper large-v3-turboOpen source
98.1%57.1%1.9%1.19×
26Qwen3-ASR-Flash
103.6%89.9%-3.6%2.07×
27OpenAI Whisper
104.6%63.8%-4.6%1.32×
28VibeVoice ASROpen source
110.1%64%-10.1%0.4×

Recorded first-party for this benchmark — audio not distributed, metrics only.