How I built a local dictation app for Mac

For the last six months I've talked to my computer more than I've typed. Prompts for Cursor and Claude Code, Slack replies, email. Apple's dictation mangled technical terms, and the cloud tools capped my words or sent my voice to a server. I wanted to hold a key, talk, let go and see the text where my cursor is, with everything running on my own Mac. Here's what turned out to be hard.

The loop

FnFlow is a native Swift and SwiftUI menu-bar app, no Electron. When you hold the dictation key it records from the microphone. When you let go the audio goes to Whisper on the same Mac, the text is pasted where your cursor was, and the app reads the field back to check that it landed. The recording and the text go to a local history. Almost every one of those steps had a trap.

The fn key

Nobody presses fn while typing, which makes it a good push-to-talk key. The catch is that macOS gives 🌐/fn its own jobs: emoji, input switching or Apple's dictation. The first versions sent people to Keyboard settings to set it to Do Nothing. Now the app reads that setting itself and uses right ⌘ when 🌐 is busy, so nobody has to leave the setup window.

Accidental taps were the surprise. 70 of my first 1,295 recordings failed, and nearly all were shorter than 1.3 seconds: I'd brushed fn on the way to another key. Recordings under 0.6 s now skip recognition entirely. I checked that threshold against the same history, and none of the good dictations was that short.

Shipping Whisper inside the app

The first builds called whisper.cpp from Homebrew, which is fine for me and useless for anyone else. Now whisper.cpp 1.9.4 is built from a pinned commit, statically, with the Metal shaders embedded. That gives two binaries of about 5 MB each with only system dependencies. They're built for armv8.4-a+fp16 rather than native, because a build tuned for my M4 Pro wouldn't start on an M1.

The very first run of a new build takes about 15 to 17 seconds while Metal compiles its pipelines, and the first dictation after install looked frozen. The app now warms the model in the background a second after launch, while you're still granting permissions.

Loading the model while you talk

Running whisper-cli for every sentence means reading the model from disk every time. For a 7.6-second sentence that was 2.26 s from releasing the key to text. So a local whisper-server on 127.0.0.1 starts the moment you press the key and loads the model while you're still speaking. The same sentence now takes 1.17 to 1.19 s on an M4 Pro, with identical text. After three idle minutes the server shuts down to give the memory back.

A model in the installer

I wanted dictation to work right after install, offline. That means shipping a model, and the full large-v3-turbo is 1.6 GB. I ran 32 of my own dictations through smaller models: small q5_1 (190 MB) held up, while base and tiny turned Russian into nonsense. So the 187 MB installer carries small q5_1, and the 8-bit large-v3-turbo (874 MB) downloads in the background on Wi-Fi. On the same 32 dictations it gave the same text as the full model. The benchmark has the details.

Pasting into Electron apps

Most dictation apps put the text on the clipboard, press ⌘V and hope. That works in native apps. Cursor, Slack, VS Code, Claude and Obsidian are Electron, and Electron hides its text fields from the Accessibility API until something asks it to expose them. Until then you can't find the field or check what's in it.

FnFlow recognizes Electron and Chromium apps by the framework in their bundle and turns on the documented AXManualAccessibility attribute once per process. It then follows the focus chain and checks that the app, window and field are the ones you started in. After pasting it reads the field back, and the pill says Pasted, Paste sent or Saved · ⌘V. Password fields are left alone, and your clipboard comes back unless you copied something new meanwhile.

Whisper's phantoms

Whisper learned from subtitles, and on silence it sometimes hears lines that were never said: “Thank you.”, “To be continued...” or a subtitle editor's credit. That happened in 33 of my 1,295 recordings. A filter removes a phantom only when it stands as its own sentence, so “Tell him thank you” stays. Across all 1,295 recordings it changed exactly those 33.

Your own words

Project names are what every model gets wrong. Whisper accepts a prompt with expected terms, so FnFlow has a vocabulary of up to 60 entries plus whole-word replacements. On a test sentence the prompt turned “update Zapot” into “update Saypad”.

What's still missing

It needs Apple Silicon and macOS 15 or later. I've tested English and Russian the most. It types exactly what you said, without AI rewriting, which some people will miss. There's no iPhone or Windows version.

Try it on your own jargon.Free · macOS 15 or later · Apple Silicon · 187 MB
Download FnFlow

Questions

Is FnFlow open source?

Not at the moment. whisper.cpp, which does the recognition, is open source (MIT), and FnFlow builds it from a pinned commit.

Who wrote the code?

Mostly AI agents, first Codex and then Claude. I acted as product owner and CTO: set the tasks, made the calls, caught the bugs and used every build daily. The numbers in this post come from measurements the agents ran on my Mac at my request.