whisper.cpp on Apple Silicon: which large-v3-turbo quantization to ship
FnFlow runs Whisper locally through whisper.cpp, so model choice is a product decision: download size against accuracy and latency. On 8 October 2026 we ran 32 real dictations through four ggml models on an M4 Pro. The short version is that q8_0 is a free lunch and q5_0 isn't.
Setup
- 14-inch MacBook Pro with M4 Pro, macOS 26. whisper.cpp 1.9.4 built statically with Metal and the shaders embedded.
- Models: ggml large-v3-turbo in f16 (1,625 MB), q8_0 (874 MB) and q5_0 (574 MB), plus small q5_1 (190 MB).
- Data: 32 real dictations from daily use, 5.4 to 18.1 s each, 339 s of audio in total. Mostly Russian with English technical terms, one speaker. Converted to 16 kHz mono WAV.
- Each file went through a fresh
whisper-cliprocess, so every timing includes loading the model (the cold path). One warm-up run per model compiled the Metal pipelines first and wasn't counted. - Metric: word-level Levenshtein distance to the f16 output divided by the f16 word count, after lowercasing and stripping punctuation.
Results
| Model (ggml) | Size | Word diff vs f16 | Identical | Median, cold | p90, cold | Time ÷ audio length |
|---|---|---|---|---|---|---|
| large-v3-turbo (full, f16) | 1,625 MB | reference | – | 2.14 s | 2.33 s | 0.23 |
| large-v3-turbo q8_0 | 874 MB | 0.0% | 31 of 32 | 1.98 s | 2.09 s | 0.21 |
| large-v3-turbo q5_0 | 574 MB | 2.4% | 21 of 32 | 2.21 s | 2.86 s | 0.25 |
| small q5_1 | 190 MB | 22.0% | 0 of 32 | 0.84 s | 1.02 s | 0.09 |
q8_0 matched f16 word for word in 31 of 32 files. The one difference was punctuation and case, so it shows as 0.0% after normalization. q5_0 changed at least one word in 11 files. small q5_1 never matched exactly but runs at under a tenth of real time.
What we took from it
q8_0 is the default download now. It's half the size of f16, gave the same text and was a little faster at both median and p90. Before 0.12 FnFlow downloaded the f16 file.
q5_0 saves another 300 MB, but it costs real words and the tail latency got worse. On a dictation app where people re-read every sentence, that's a bad trade.
small q5_1 is what we bundle in the installer. The first dictation works offline right after install, and that matters more than accuracy for the first minute. It's not good enough to stay on, so the app swaps to q8_0 when the download finishes.
Cold vs warm
The numbers above are cold: process start, model load, inference. In the app a local whisper-server starts the moment you press the key, so the model loads while you're still talking and stays resident for three minutes after the last dictation. With that, a 7.6-second sentence comes back in 1.17 to 1.19 s after key release, against 2.26 s for a cold whisper-cli run on the same audio.
The one slow case is the very first run of a new build. Metal compiles its pipelines, which took about 15 to 17 seconds on this machine. FnFlow does that in the background right after launch so the first real dictation doesn't wait.
Reproduce it
afconvert -f WAVE -d LEI16@16000 -c 1 input.caf input.wav
whisper-cli -m ggml-large-v3-turbo-q8_0.bin -f input.wav -l auto -nt -np
whisper-server -m ggml-large-v3-turbo-q8_0.bin -l auto -t 4 -fa --host 127.0.0.1 --port 8080
Models are on Hugging Face in ggerganov/whisper.cpp. A plain-language version of this test is on Saypad. Your numbers will move with the chip, the language and how much you mumble, so treat ours as one data point.
Questions
Why is q5_0 slower than q8_0 here?
We didn't profile it, so this is a guess: on Metal the extra dequantization work for 5-bit weights seems to cost more than the smaller read saves at this model size. The p90 gap (2.86 s vs 2.09 s) was bigger than the median gap.
Is word diff against f16 the same as WER?
No. It's word-level Levenshtein distance against the f16 model's output, divided by the reference word count. There's no human transcript, so it measures what quantization costs you, not absolute accuracy.
What about tiny and base?
An earlier run on 32 recordings (24 Russian, 8 English) put tiny at 52% / 65% word diff against large-v3-turbo and base at 35% / 65% (Russian / English). Both produced nonsense on Russian. small q5_1 was 24% / 41% on that set and is the smallest model we'd ship.
Does flash attention or thread count matter?
The benchmark used whisper-cli defaults. FnFlow's warm server runs with -fa and -t 4. We haven't measured those flags separately yet.
Can I use these models with FnFlow?
FnFlow ships small q5_1 and downloads large-v3-turbo q8_0 in the background. The full f16 file is still accepted if it's already on disk.