モンゴル語 音声テキスト変換リーダーボード
モンゴル語の音声テキスト変換プロバイダの単語誤り率・速度・価格を比較します。すべての数値は実際のモンゴル語音声で実測 — 宣伝ではありません。
使用したデータセット
各プロバイダーの実行に使用した公開音声コーパスです。ソース URL を開くと、データセットの説明・ライセンス・ダウンロードを確認できます。
- Common Voice 24 (MN)›https://huggingface.co/datasets/btsee/common-voices-24-mn
- Shunya Labs Mongolian Speech›https://huggingface.co/datasets/shunyalabs/mongolian-speech-dataset
- Common Voice 20 (MN)›https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized
- Modern Voice›このベンチマークのために自社で録音 — 音声は配布せず、指標のみ公開します。
WER vs 速度
各ドットが1プロバイダ。好ましいのは右下(誤り少・高速)。実際のモンゴル語音声で実測。
Google STTDuudlaga FlowSpeechmaticsChimegeOpenAI WhisperGPT-4o TranscribeAzure SpeechGladiaGemma 4Whisper large-v3Whisper large-v3-turboWhisper large-v2 MNSeamlessM4T v2MMS-1B-allMoonshine MNwav2vec2 XLSR-53 MNWhisper turbo MNWhisper medium MNElevenLabs Scribe v1ElevenLabs Scribe v2Qwen3-ASR-FlashDolphin smallOmniASR CTC-1BOmniASR LLM-1BVibeVoice ASRWhisper large-v3 MN
プロバイダ比較
各列を並べ替えできます。「詳細」でプロバイダごとのサンプル分解・価格・手法を確認。
| プロバイダ | WER ↑ | CER | 精度 | 速度 | レイテンシ | 価格 / 1000分 | |
|---|---|---|---|---|---|---|---|
Speechmatics | 14.4% | 7.4% | 85.6% | 0.51× | 13.8s | $8.5 | 詳細 |
Duudlaga Flow自社 | 14.6% | 7.6% | 85.4% | 1.47× | 5.2s | $19.4 | 詳細 |
Chimege | 14.6% | 6.6% | 85.4% | 1.51× | 4.7s | $11.1 | 詳細 |
Google STT | 16.9% | 7.4% | 83.1% | 1.98× | 3.7s | $16 | 詳細 |
Whisper large-v2 MN | 20.7% | 8.3% | 79.3% | 0.53× | 13.5s | $0 | 詳細 |
Azure Speech | 20.8% | 9.7% | 79.2% | 0.39× | 17.3s | $16.7 | 詳細 |
Whisper medium MN重複 | 22.5% | 8.9% | 77.5% | 2.7× | 2.7s | $0 | 詳細 |
SeamlessM4T v2 | 24% | 11.1% | 76% | 0.39× | 26.3s | $0 | 詳細 |
Whisper turbo MN重複 | 25.7% | 9.2% | 74.3% | 2.43× | 3.0s | $0 | 詳細 |
ElevenLabs Scribe v1 | 27.1% | 10.8% | 72.9% | 2.14× | 3.3s | $6.7 | 詳細 |
ElevenLabs Scribe v2 | 27.1% | 10.6% | 72.9% | 1.51× | 4.7s | $3.7 | 詳細 |
Gemma 4 | 30.2% | 14.7% | 69.8% | 1× | 7.3s | $0 | 詳細 |
OmniASR LLM-1B | 33.8% | 14.2% | 66.2% | 0.27× | 26.6s | $0 | 詳細 |
Whisper large-v3 MN重複 | 39.4% | 14.9% | 60.6% | 1.57× | 4.5s | $0 | 詳細 |
MMS-1B-all | 41.7% | 12.3% | 58.3% | 19.2× | 0.4s | $0 | 詳細 |
wav2vec2 XLSR-53 MN重複 | 45% | 16.4% | 55% | 42.32× | 0.2s | $0 | 詳細 |
Dolphin small | 49.1% | 19.2% | 50.9% | 1.73× | 4.1s | $0 | 詳細 |
OmniASR CTC-1B | 51.2% | 15.5% | 48.8% | 0.6× | 11.7s | $0 | 詳細 |
GPT-4o Transcribe | 53% | 27.6% | 47% | 2.11× | 3.7s | $6 | 詳細 |
Moonshine MN重複 | 58.8% | 45.2% | 41.2% | 8.44× | 0.8s | $0 | 詳細 |
Whisper large-v3 | 89.5% | 37.4% | 10.5% | 1.29× | 6.0s | $0 | 詳細 |
Gladia | 90.1% | 38.9% | 9.9% | 0.75× | 9.4s | $10.2 | 詳細 |
Whisper large-v3-turbo | 99% | 54.4% | 1% | 0.91× | 8.7s | $0 | 詳細 |
Qwen3-ASR-Flash | 103.7% | 86.4% | -3.7% | 2.19× | 3.2s | $2.1 | 詳細 |
OpenAI Whisper | 105.2% | 63.7% | -5.2% | 1.5× | 4.5s | $6 | 詳細 |
VibeVoice ASR | 107% | 61% | -7% | 0.53× | 12.1s | $0 | 詳細 |
Gemini Flash | ベンチマーク近日公開 | ||||||
すべて実測のモンゴル語データセットに基づきます。速度は×実時間(大きいほど高速)。WERは低いほど良い。
重複— このバッジが付いたモデルは、ベンチマークのコーパスと重複するデータで学習しています(モンゴル語 Common Voice の 173 サンプル中 113 件、または Shunya Labs の公開コーパス)。重複するコーパスでのスコアは記憶によって高く出ます。これらのモデルは詳細ページのデータセット別の表、とくに学習に使われていないコーパスで判断してください。
Duudlaga Voice Set
92 samples · 10.2 min · 28 systemsA second corpus recorded first-party for this benchmark, covering what the public Mongolian datasets barely contain: modern loanwords, English/Mongolian code-switching, numbers, dates, and commands. Every system is measured on the same 92 recordings.
| プロバイダ | WER | CER | 精度 | 速度 |
|---|---|---|---|---|
1Duudlaga Flow自社 | 9.5% | 5.3% | 90.5% | 1.27× |
2Google STT | 11.7% | 5.5% | 88.3% | 2.28× |
3Chimege | 14.5% | 6.8% | 85.5% | 1.48× |
4Speechmatics | 15.3% | 7.6% | 84.7% | 0.5× |
5Gemini 3.5 Flash | 19.5% | 11.4% | 80.5% | 0.06× |
6SeamlessM4T v2オープンソース | 22.9% | 10.4% | 77.1% | 4.32× |
7Whisper large-v2 MNオープンソース | 23.2% | 10.9% | 76.8% | 0.55× |
8ElevenLabs Scribe v1 | 23.4% | 9.3% | 76.6% | 2.09× |
9ElevenLabs Scribe v2 | 24.6% | 9.8% | 75.4% | 1.47× |
10Azure Speech | 26.1% | 10.9% | 73.9% | 0.32× |
11Gemma 4オープンソース | 27.7% | 13.2% | 72.3% | 1.04× |
12OmniASR LLM-1Bオープンソース | 34.5% | 14.5% | 65.5% | 0.27× |
13Whisper medium MNオープンソース重複 | 35.1% | 14.8% | 64.9% | 2.71× |
14Whisper turbo MNオープンソース重複 | 37.6% | 14.1% | 62.4% | 2.59× |
15Whisper large-v3 MNオープンソース重複 | 41% | 17.1% | 59% | 1.58× |
16GPT-4o Transcribe | 44.3% | 23.7% | 55.7% | 2.62× |
17MMS-1B-allオープンソース | 45.3% | 15.1% | 54.7% | 20.52× |
18Dolphin smallオープンソース | 46.1% | 19% | 53.9% | 1.69× |
19W2v-BERT 2.0 MNオープンソース重複 | 46.3% | 18.2% | 53.7% | 28.12× |
20OmniASR CTC-1Bオープンソース | 54.9% | 18.9% | 45.1% | 0.59× |
21wav2vec2 XLSR-53 MNオープンソース重複 | 56.3% | 23.9% | 43.7% | 45.7× |
22Whisper large-v3オープンソース | 87.3% | 36.6% | 12.7% | 1.55× |
23Gladia | 88.1% | 39.9% | 11.9% | 0.74× |
24Moonshine MNオープンソース重複 | 96.2% | 75.7% | 3.8% | 8.36× |
25Whisper large-v3-turboオープンソース | 98.1% | 57.1% | 1.9% | 1.19× |
26Qwen3-ASR-Flash | 103.6% | 89.9% | -3.6% | 2.07× |
27OpenAI Whisper | 104.6% | 63.8% | -4.6% | 1.32× |
28VibeVoice ASRオープンソース | 110.1% | 64% | -10.1% | 0.4× |
このベンチマークのために自社で録音 — 音声は配布せず、指標のみ公開します。