廣東話語音辨識基準Cantonese ASR benchmark
最後更新:2026 年 9 月 11 日 · 6,933 句 · 8 小時 21 分鐘語音 · 六個範疇 Last updated 11 September 2026 · 6,933 utterances · 8 h 21 min of audio · six domains
廣東話語音辨識冇一個公開嘅排行榜。你想知邊個模型準啲, 你要自己行一次。我哋行咗,所以擺出嚟。
入面有一欄係全世界都冇人報過嘅:中英夾雜。 Google、Azure、OpenAI、ElevenLabs、Apple——冇一間公佈過中英夾雜嘅錯誤率。 但係喺香港,一句入面轉幾轉語言先係常態,唔係例外。所以嗰欄先係真正決定 一個 app 用唔用得嘅數字。
下面全部數字都係量度返嚟嘅,方法同埋佢行嘅數據集都係公開嘅, 你可以自己行一次去質疑我哋。包括我哋輸嘅地方。
There is no public leaderboard for Cantonese speech recognition. If you want to know which model is more accurate, you have to run it yourself. We did, so here it is.
One column in it has never been published by anyone: code-switching. Google, Azure, OpenAI, ElevenLabs and Apple all decline to report a Chinese-English code-switching error rate. Yet in Hong Kong, switching language several times inside one sentence is the norm rather than the exception — which makes that column the number that actually decides whether an app is usable.
Every figure below is measured. The method, and the dataset it runs on, are public, so you can run it yourself and argue with us. That includes the places we lose.
15.45%
Cantalkie 出廠模型喺中英夾雜語音上面嘅字元錯誤率, 喺呢個公開基準嘅 Yue & Eng 部分度量。同一組音頻, 第二個常見嘅開源模型係 20.49% 同 20.93%。
Character error rate of the model Cantalkie ships, on code-switched speech, measured on the Yue & Eng subset of this public benchmark. On the same audio the two other common open models score 20.49% and 20.93%.
六個範疇嘅完整結果The full result, six domains
字元錯誤率(CER),越低越好。全部喺同一組音頻、同一套評分程式行出嚟。
Character error rate (CER); lower is better. All rows on the same audio, scored by the same code.
hon9kon9ize/cantonese_asr_eval,6,933 句。Cantalkie 嗰欄係淨 ASR,所有修正層關咗。藍色=嗰行最低嘅數字,唔理係邊個。 hon9kon9ize/cantonese_asr_eval, 6,933 utterances. The Cantalkie column is bare ASR, every correction layer switched off. Blue marks the lowest figure in each row, whoever it belongs to.
← 個表可以左右拉 →← the table scrolls sideways →
| 範疇Domain | Qwen3-ASR-1.7B 8-bit,自動語言 (Cantalkie 出廠設定), auto language (as Cantalkie ships) |
Qwen3-ASR-1.7B bf16 |
Qwen3-ASR 0.6B |
SenseVoice Small |
|---|---|---|---|---|
| Mixed | 6.39% | 6.21% | 8.69% | 7.69% |
| Daily Use | 4.14% | 4.22% | 6.12% | 5.64% |
| Commands | 4.85% | 4.89% | 6.83% | 7.45% |
| Yue & Eng | 15.45% | 15.37% | 20.93% | 20.49% |
| Storytelling | 11.79% | 11.79% | 14.31% | 14.52% |
| Synthetic | 6.87% | 7.04% | 9.65% | 8.50% |
兩個要講清楚嘅地方。 一,Cantalkie 嗰欄係淨 ASR —— 粵拼對照、詞庫加權、繁簡修正、語言模型執稿全部冇行。呢個唔係手軟, 係因為一加咗呢啲層,呢一欄就冇得同人比。真實使用嘅輸出好過呢個數字。 二,bf16 嗰欄喺 Mixed 同 Yue & Eng 係贏 8-bit 嘅 —— 差 0.18 同 0.08 個百分點。 我哋照登。
Two things to be clear about. First, the Cantalkie column is bare ASR: no Jyutping lookup, no vocabulary biasing, no script correction, no language-model clean-up. That is not modesty — adding those layers would make the column incomparable with everyone else's. What a user actually sees is better than this number. Second, bf16 beats our shipped 8-bit on Mixed and on Yue & Eng, by 0.18 and 0.08 percentage points. We are publishing that too.
四個量度返嚟嘅發現Four measured findings
一、8-bit 量化係免費嘅1. 8-bit quantisation is free
淨睇精度:同一個模型、同樣指定語言,bf16 對 8-bit, 平均絕對差 0.06 個百分點,最大 0.12。三個範疇量化之後仲好啲, 兩個差啲。呢個係雜訊,唔係損失。之前冇人量度過。
上面個表嗰兩欄之差係 0.09 個百分點, 數字大啲,因為嗰個比較同時變咗兩樣嘢 —— 精度,同埋語言模式 (我哋出廠係自動,佢哋嗰行係指定咗廣東話)。所以擺喺呢度講清楚, 而唔係將兩個數字撈埋一齊。
Precision alone: same model, both with the language forced, bf16 against 8-bit — mean absolute difference 0.06 percentage points, largest single move 0.12. Three domains got better under quantisation, two got worse. That is noise, not loss. Nobody had measured it before.
The gap between those two columns in the table above is 0.09 percentage points, slightly larger, because that comparison moves two things at once — precision and language mode (we ship auto; their row is forced). Which is why both numbers are stated here rather than blended into one.
二、指定語言唔使付出代價,仲快三成2. Forcing the language costs nothing and saves 30% of the time
指定 language=Cantonese 對自動偵測,平均絕對差
0.03 個百分點;喺中英夾雜嗰組,兩者差 0.01。
但係指定語言快咗大約 三成 解碼時間。
呢個推翻咗我哋自己講咗幾個星期嘅嘢。我哋一直同人講「指定語言會搞砸中英夾雜」, 喺句子長度嘅音頻上面,呢個講法唔成立。短過一秒嘅片段仲係另一回事, 嗰個問題未解決。
出廠設定仍然係自動,冇因為呢個發現就改。 指定語言嗰次行程嘅逐項數字未寫入檔案,其中一格喺筆記入面記錄得唔一致, 所以呢版唔會登嗰一欄 —— 寧願少登一欄,都唔可以登一個追唔返原始紀錄嘅數字。
Forcing language=Cantonese against auto-detection: mean absolute
difference 0.03 percentage points, and on the code-switching subset the two
agree to 0.01. Forcing is about 30% faster to decode.
This overturned something we had been saying for weeks. We told people that forcing a language breaks code-switching. At sentence length, that claim is not supported. Sub-second clips remain a separate, unsolved question.
The shipped default is still auto; this finding did not change it. The per-domain figures from the forced run were never written to a results file and one cell is recorded inconsistently in the working notes, so that column is not printed here. Better a missing column than a number that cannot be traced back to a stored run.
三、有啲「錯誤率」其實係大階字3. Some of that "error rate" is just block capitals
我哋 Android 版用嘅 SenseVoice checkpoint,喺中英夾雜嗰組原始
CER 係 63.43%。睇落災難級。但係佢將英文全部寫成
YOU KNOW WHAT I MEAN,而參考答案係細階。
淨係改返大細階,個數字就由 63.43% 跌到 22.75%,
其他五個範疇郁少過 0.05。
所以任何用原始 CER 去比較中英夾雜嘅講法,可能只係喺度量緊大細階。 唔好引用 63% 嗰個數。
The SenseVoice checkpoint our Android build uses scores a raw CER of
63.43% on the code-switching subset. That looks catastrophic. But it writes
English as YOU KNOW WHAT I MEAN while the references are lower-case. Normalising
case alone moves that cell from 63.43% to 22.75%, and every other domain by
less than 0.05.
So any raw-CER comparison of code-switching may be measuring capitalisation. Do not quote the 63% figure — including against us.
四、CER 睇唔到「係」定「系」4. CER cannot see the difference between 係 and 系
香港人讀到「唔係」定「唔系」, 係睇得出嘅分別。但係評分程式行 t2s 之後先評分,兩個字會被收埋成同一個, 所以呢個分別對 CER 嚟講價值係零。
同一次行程入面:修正層關咗,「係」寫啱嘅比率係 12.0%; 開返,83.0%;6,933 句入面 1,268 句(18.3%)改咗。 而六個範疇嘅 CER 到小數點後兩位完全冇郁過。
呢個就係點解唔可以淨係睇一個數字去排優先次序。
A Hong Kong reader can see the difference between 唔係 and 唔系 immediately. The scoring script runs a Traditional-to-Simplified pass before scoring, which collapses the two — so that distinction is worth exactly zero to CER.
In the same run: with the correction layer off, 係 is written correctly 12.0% of the time; with it on, 83.0%. It changes 1,268 of 6,933 utterances (18.3%). CER across all six domains is identical to two decimal places.
Which is why no one should rank this kind of work by a single number, ours included.
同 Apple 內建聽寫比Against Apple's built-in dictation
首先講清楚:Apple 係支援廣東話嘅。
zh-HK 出繁體字,口語廣東話處理得好好。
任何話「Apple 唔識廣東話」嘅講法都係錯嘅,包括競爭對手講。
First, plainly: Apple does support Cantonese. Its
zh-HK locale outputs Traditional characters and handles colloquial Cantonese
properly. Anyone telling you Apple cannot do Cantonese is wrong — competitors included.
一段 29 分半鐘、兩個人講嘢嘅真實廣東話工作會議錄音。M5 Pro,macOS 26.5。 One real 29½-minute two-speaker Cantonese work call. M5 Pro, macOS 26.5.
← 個表可以左右拉 →← the table scrolls sideways →
Apple zh-HK |
Cantalkie | |
|---|---|---|
| 處理時間Wall clock | 11.8 s | 153 s |
| 轉出字數Characters transcribed | 7,442 | 10,401 |
| 標點符號Punctuation marks | 96 | 1,097 |
| 保留到嘅英文段落English runs kept | 199 | 276 |
Apple 快十三倍,呢個我哋輸得乾脆。 但係佢少轉咗大約 28% 嘅字 —— 佢成段成段咁漏咗較細聲嗰位講者。 單人聽寫冇問題,開會就唔夠用。
中英夾雜先係分野。有一段 45 秒嘅錄音,12 個英文詞 Apple 啱咗大概 4 個
(「key points」變 keypound,仲有 documert、Cludie、
unabe);Cantalkie 15 個啱 14 個。
另一處「我已經收咗 offer」,Apple 聽成「我已經收到 oven」。
誠實講輸嘅地方:速度輸到冇得拗;純廣東話準確度係打和 —— Apple 有一次寫「問吓」,比我哋嘅「問下」仲地道。
呢個比較嘅限制。一段錄音、一個環境, 冇做人手標準答案,用字數同標點數做覆蓋率嘅代替指標。 呢個係方向性嘅證據,唔係基準。而且測嘅係開發者 API,唔係用家撳掣嗰個聽寫。 我哋照咁講,因為呢啲限制唔講就變咗誤導。
Apple is about thirteen times faster. We lose that outright. But it transcribed roughly 28% fewer characters, dropping whole stretches of the quieter speaker. Fine for single-speaker dictation; not enough for a meeting.
Code-switching is where they separate. In one 45-second passage Apple got about
4 of 12 English terms right — keypound for "key points", plus documert,
Cludie and unabe; Cantalkie got 14 of 15. Elsewhere Apple heard
「我已經收咗 offer」 as 「我已經收到 oven」.
The honest losses: speed, decisively; and pure-Cantonese accuracy is level — Apple once wrote 問吓, which is more idiomatic than our 問下.
Limits of this comparison. One recording, one acoustic environment, no hand-made ground-truth transcript; character and punctuation counts stand in for coverage. This is directional evidence, not a benchmark. It also tested the developer API rather than the dictation users press a key for. We say so because leaving it out would make the comparison misleading.
方法Method
- 數據集:hon9kon9ize/cantonese_asr_eval —— 呢個社群本身已經用緊嘅基準,唔係我哋自己砌嘅。六個子集,6,933 句,8 小時 21 分鐘。
- 硬件:M5 Pro。實時比率 0.040–0.056(指定語言)/0.057–0.079(自動偵測)。
- Cantalkie 嗰欄:淨 ASR。冇粵拼表、冇詞庫加權、冇 OpenCC、冇語言模型執稿。
- 評分:我哋重寫咗
eval.py(冇用evaluate/jiwer), 每個子集用返佢哋原本嘅種子同抽樣數目。 - 其他模型嗰幾欄:直接攞自該倉庫已經 commit 咗嘅
results/目錄, 唔係我哋自己行嘅。 - 我哋嘅 code 已經開返 PR 畀佢哋,等我哋嗰兩行有得畀人查證。
- Dataset: hon9kon9ize/cantonese_asr_eval — the benchmark this community already uses, not one we invented for ourselves. Six subsets, 6,933 utterances, 8 h 21 min.
- Hardware: M5 Pro. Real-time factor 0.040–0.056 forced, 0.057–0.079 auto.
- The Cantalkie column: bare ASR. No Jyutping table, no vocabulary biasing, no OpenCC, no language-model rewrite.
- Scoring: we reimplemented
eval.pywithoutevaluate/jiwer, replaying each subset with its original seeds and sample sizes. - The other model columns are taken from that repository's own committed
results/directory, not re-run by us. - Our model class has been offered back as a pull request, so our two rows can be checked rather than taken on trust.
一個我哋想搞清楚嘅落差A discrepancy we would like resolved
我哋嘅方法喺 MiMo 同 FireRedASR 六個子集全部重現到,
喺 Mixed/Daily Use/Commands 三個子集所有模型都重現到。
但係 SenseVoice 同兩個 Whisper 喺另外三個子集重現唔到 —— 最大嗰格係
SenseVoice · Yue & Eng:README 寫 9.05%,佢哋自己 results/
入面係 20.49%。七個模型嘅 expected 清單雜湊值完全一樣,
即係數據冇變過,最可能係嗰幾個模型重行過而 README 冇更新。
我哋擺出嘅係 results/ 嗰個數。呢點已經喺佢哋 Discord 同
PR 描述入面提出咗,語氣係「我哋可能搞錯」。我哋冇改動過佢哋任何數字。
Our method reproduces MiMo and FireRedASR on all six subsets, and every model on
Mixed, Daily Use and Commands. It does not reproduce SenseVoice or either Whisper on the other
three. The worst cell: SenseVoice · Yue & Eng — the README says 9.05%, that
repository's own results/ directory says 20.49%. The expected
lists hash identically across all seven models, so the data has not changed; most likely those
models were re-run and the README was never updated.
We publish the results/ figure. This has been raised in their
Discord and in the pull request description, with room for us to be the ones who are wrong. We
have altered none of their numbers.
點樣自己查證How to check this yourself
Clone 上面個基準倉庫,行返佢自己嘅 eval.py。
你會攞到 results/ 嗰幾欄。想查我哋嗰兩行,等個 PR 合併,
或者電郵 support@cantalkie.com,
我哋將原始輸出寄畀你。
Clone the benchmark repository above and run its own eval.py; that
gives you the results/ columns. To check our two rows, wait for the pull request to
merge, or email support@cantalkie.com and we will
send you the raw output.