Mongolian Speech-to-Text Leaderboard
Qwen3-ASR-FlashOperating point: qwen3-asr-flash

Qwen3-ASR-Flash — Mongolian Speech-to-Text Benchmark

Real results from the Qwen3-ASR-Flash Batch API (enhanced operating point, language auto (mn rejected)) evaluated across 4 Mongolian datasets and 265 audio samples. Every number below is measured — not marketing. WER, speed, and pricing are shown as-is, brutally honest.

Language:AUTO (MN REJECTED)Diarization:noneSamples:265Run:Aug 24, 2026, 3:13 PMEndpoint:https://dashscope-intl.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
067100-3.7%
Accuracy

265/265 samples transcribed · 100% success rate

067100103.7%
Word Error Rate

Lower is better · across 265 samples

Headline metrics
Character Error Rate86.4%
Avg speed factor2.19×realtime multiple
Total speed factor2.19×
Avg latency / sample3.2s
Total audio processed1849.3s30.8 min

Pricing

Qwen3-ASR-Flash list pricing for batch transcription. No discounts, no negotiated rates applied — the raw per-minute rate.

Per 1k minutes
$2.1
batch
Per minute
$0.0021
effective
Per 1k min (these 30.8 min)
$3.88
would cost
Open source?
proprietary

Pricing source: Qwen3-ASR-Flash public pricing. Duudlaga Flow is shown for context only — this page isolates Qwen3-ASR-Flash so the number is not padded by our own product.

WER & CER by dataset

Word and character error rates per dataset. Lower is better — and these are the real Qwen3-ASR-Flash numbers, which are weak on Common Voice 24.

Common Voice 24 (MN)
Source dataset →
WER
105.2%
CER
90%
Shunya Labs Mongolian Speech
Source dataset →
WER
102.3%
CER
75.2%
Common Voice 20 (MN)
Source dataset →
WER
104.5%
CER
90.4%
Modern Voice
WER
103.2%
CER
89.1%
WERCER

Dataset summary

Aggregate accuracy, speed, and timing for each dataset.

DatasetSourceSamplesSuccessWERCERAccuracySpeedAudio (s)Proc (s)
Common Voice 24 (MN)Hugging Face →5959/59105.2%90%-5.2%2.81×330.6s117.7s
Shunya Labs Mongolian SpeechHugging Face →6060/60102.3%75.2%-2.3%1.94×645.1s332.9s
Common Voice 20 (MN)Hugging Face →5454/54104.5%90.4%-4.5%2.71×269.8s99.5s
Modern Voice9292/92103.2%89.1%-3.2%2.04×603.8s295.3s

WER vs speed — per sample

Each dot is one audio sample. The sweet spot is the bottom-left (low error, fast). Qwen3-ASR-Flash sits high on error for many Common Voice 24 samples.

Per-sample results

Ground truth shown verbatim in the Expected column. The result column highlights only the words Qwen3-ASR-Flash got wrong, in red — no strikethrough/swap gymnastics, just the mistakes.

265 rows
redwrong word in the resultplaincorrectly transcribedExpected column shown verbatim as ground truth
#AudioSampleDatasetExpected (ground truth)Qwen3-ASR-Flash resultWERCERI/D/S
1btsee_0001Common Voice 24 (MN)Гэхдээ амьсгал хураахаасаа өмнө танд мэдэж байгаагаа хэлье.كشدا مسخرة خاصة ومن تندم يدي جاغا هي لي.112.5%87.9%1/0/8
2btsee_0002Common Voice 24 (MN)Надад заяасан аз жаргал гэдэг ердөө гуравхан сарын хугацаатай байсан гэж үү?나타자야상아처럼깨끗이剃도그러곤살에혹자떼버리는게좋.100.0%100.0%0/11/1
3btsee_0003Common Voice 24 (MN)Одоо бид өөрсдөө өвчин эмгэгээсээ салахыг хичээцгээе.أعطى بيدروس دوتشين فيكيسي سفيرًا لتشيكيا.100.0%92.3%0/1/6
4btsee_0004Common Voice 24 (MN)Би бол голдуу хээрээр гэр, хэцээр дэр хийж явдаг хүн.Við varst kalt og hér er gott, ég þyrfti þér ekki segja það kom.150.0%111.8%5/0/10
5btsee_0005Common Voice 24 (MN)Хан хурмаст уурлаж, Болдоггүй Бор өвгөнийг хор луугаараа ниргүүлэхээр явуулжээ.خانفرم استورت سپورت کویبرا ونیو خیرت ثووگارنیرو شریمچی.100.0%92.2%0/2/8
6btsee_0006Common Voice 24 (MN)Алив наашаа ороод ир гээд гэртээ оров.عجبنه شعرة كتكشت تعرف.100.0%91.9%0/3/4
7btsee_0007Common Voice 24 (MN)Өө өндөр дээдэс таны тухайд би баталж чадахгүй.أو أنظر تلك التنيطات ببطء شتى.100.0%89.1%0/2/6
8btsee_0008Common Voice 24 (MN)Харин гурав дахь удаагаас эхлэн хүмүүсийг сонирхож эхлэв.Herhangi rüptükle takas etmek umutsuzca zor çıktı.100.0%92.9%0/1/7
9btsee_0009Common Voice 24 (MN)Та нар очингуутаа шөл л өгч үз.Таны орчингод шүдлэл тохуч.100.0%60.0%0/3/4
10btsee_0010Common Voice 24 (MN)Ерөөсөө литр үйлдвэрээсээ салаагүй явсан юм чинь.이로써 롤드컵에서 사타구니 없앤 첫.100.0%91.7%0/2/5
Page 1 of 27

Methodology

How these numbers were produced.

Provider: Qwen3-ASR-Flash (Alibaba qwen3-asr-flash hosted transcription with automatic language detection — Mongolian is not a supported language (the API rejects the 'mn' code).).

Endpoint https://dashscope-intl.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation. Language auto (mn rejected). Operating point qwen3-asr-flash. Diarization none.

Datasets: Common Voice 24 (MN), Shunya Labs Mongolian Speech, Common Voice 20 (MN), Modern Voice — 265 samples, 1849.3s of audio total.

Dataset source URLs:

Metrics: WER and CER are computed with a standard word/character Levenshtein alignment, normalized for case and punctuation. Accuracy = 100 − WER. Speed factor = audio duration ÷ processing time (× realtime). All requests are real Qwen3-ASR-Flash Batch API calls, not cached or simulated.

Diff highlighting: The result column aligns to the ground truth and colors every substitution and insertion red. Deletions (words missing from the result) are not shown in the result column — the Expected column already holds the full ground truth as-is.

Generated by the Qwen3-ASR-Flash Benchmark Runner · Qwen3-ASR-Flash Batch API v2 · run Aug 24, 2026, 3:13 PM