whisper.cpp on Apple Silicon: which large-v3-turbo quantization to ship

FnFlow runs Whisper locally through whisper.cpp, so model choice is a product decision: download size against accuracy and latency. On 8 October 2026 we ran 32 real dictations through four ggml models on an M4 Pro. The short version is that q8_0 is a free lunch and q5_0 isn't.

Setup

Results

Model (ggml)SizeWord diff vs f16IdenticalMedian, coldp90, coldTime ÷ audio length
large-v3-turbo (full, f16)1,625 MBreference–2.14 s2.33 s0.23
large-v3-turbo q8_0874 MB0.0%31 of 321.98 s2.09 s0.21
large-v3-turbo q5_0574 MB2.4%21 of 322.21 s2.86 s0.25
small q5_1190 MB22.0%0 of 320.84 s1.02 s0.09

q8_0 matched f16 word for word in 31 of 32 files. The one difference was punctuation and case, so it shows as 0.0% after normalization. q5_0 changed at least one word in 11 files. small q5_1 never matched exactly but runs at under a tenth of real time.

What we took from it

q8_0 is the default download now. It's half the size of f16, gave the same text and was a little faster at both median and p90. Before 0.12 FnFlow downloaded the f16 file.

q5_0 saves another 300 MB, but it costs real words and the tail latency got worse. On a dictation app where people re-read every sentence, that's a bad trade.

small q5_1 is what we bundle in the installer. The first dictation works offline right after install, and that matters more than accuracy for the first minute. It's not good enough to stay on, so the app swaps to q8_0 when the download finishes.

Cold vs warm

The numbers above are cold: process start, model load, inference. In the app a local whisper-server starts the moment you press the key, so the model loads while you're still talking and stays resident for three minutes after the last dictation. With that, a 7.6-second sentence comes back in 1.17 to 1.19 s after key release, against 2.26 s for a cold whisper-cli run on the same audio.

The one slow case is the very first run of a new build. Metal compiles its pipelines, which took about 15 to 17 seconds on this machine. FnFlow does that in the background right after launch so the first real dictation doesn't wait.

Reproduce it

afconvert -f WAVE -d LEI16@16000 -c 1 input.caf input.wav
whisper-cli -m ggml-large-v3-turbo-q8_0.bin -f input.wav -l auto -nt -np
whisper-server -m ggml-large-v3-turbo-q8_0.bin -l auto -t 4 -fa --host 127.0.0.1 --port 8080

Models are on Hugging Face in ggerganov/whisper.cpp. A plain-language version of this test is on Saypad. Your numbers will move with the chip, the language and how much you mumble, so treat ours as one data point.

Local Whisper, warm and ready.Free · macOS 15 or later · Apple Silicon · 187 MB
Download FnFlow

Questions

Why is q5_0 slower than q8_0 here?

We didn't profile it, so this is a guess: on Metal the extra dequantization work for 5-bit weights seems to cost more than the smaller read saves at this model size. The p90 gap (2.86 s vs 2.09 s) was bigger than the median gap.

Is word diff against f16 the same as WER?

No. It's word-level Levenshtein distance against the f16 model's output, divided by the reference word count. There's no human transcript, so it measures what quantization costs you, not absolute accuracy.

What about tiny and base?

An earlier run on 32 recordings (24 Russian, 8 English) put tiny at 52% / 65% word diff against large-v3-turbo and base at 35% / 65% (Russian / English). Both produced nonsense on Russian. small q5_1 was 24% / 41% on that set and is the smallest model we'd ship.

Does flash attention or thread count matter?

The benchmark used whisper-cli defaults. FnFlow's warm server runs with -fa and -t 4. We haven't measured those flags separately yet.

Can I use these models with FnFlow?

FnFlow ships small q5_1 and downloads large-v3-turbo q8_0 in the background. The full f16 file is still accepted if it's already on disk.