How it works

The engineering, not the pitch.

Local mode has two on-device ASR engines; optional Plainsay Cloud and bring-your-own-provider modes handle speech remotely. A best-effort cleanup pass never gets to lose a dictation to a bad network call. Here are the real numbers and actual tradeoffs.

Local mode: two fully on-device engines

Local mode ships two speech engines through Core ML, and you pick which one runs. Recorded audio stays on your Mac in this mode:

  • Whisper, via WhisperKit — OpenAI's model family, broad language coverage, running as a Core ML model on the Neural Engine.
  • NVIDIA Parakeet TDT 0.6B v3, via FluidAudio — faster and, in our disclosed 50-clip English benchmark, measurably more accurate; the model also supports Polish.

The Parakeet integration has a couple of specific choices worth stating plainly:

  • The encoder uses int8 precision to reduce model size and inference cost. Plainsay calls FluidAudio's download step, resolves the directory that actually contains the fetched subset, verifies every present pinned file, and only then calls the load step. This benchmark does not compare int8 with fp16, so it makes no claim about the accuracy difference between them.
  • Mel-context carry-over is disabled (ASRConfig(melChunkContext: false)) for multilingual v3 recordings. FluidAudio's pinned v0.15.6 long-transcription notes document this path as protection against wrong-language drift at chunk boundaries.
  • Decoder state is fresh for every single utterance — a new TdtDecoderState per transcribe() call, deliberately never reused across dictations, so nothing you said in one dictation can leak linguistic context into the next.

Which model to default to isn't hardcoded, either. OnDeviceModel.recommended(for:) takes the languages you actually speak (set once in the Setup Assistant) and checks them against FluidAudio's own Language.allCases — Parakeet's real, current language coverage, not a table we maintain by hand and let go stale. Speak a language outside that list and the recommendation falls back to Whisper automatically.

The main spoken language matters. Providers and models that accept only one language hint receive it, while the complete list tells Plainsay which languages are expected. When Whisper detects a language outside that list, or the transcript contains narrow character evidence of an unlisted language, Plainsay retries in the main language. Parakeet exposes neither a detected language nor a forced-language retry, so the same high-confidence evidence is recorded in diagnostics instead of being presented as a correction the engine cannot make.

Vocabulary corrections run locally after transcription, including with Parakeet and when Polishing is off. Matching uses conservative edit thresholds, scores every candidate, and chooses the closest one rather than letting the first merely acceptable term rewrite ordinary words.

Polishing: best-effort, with a raw-text fallback

Raw ASR output is grammatically honest but reads like speech: filler words, false starts, no punctuation. The optional Polishing pass turns that into written text via an LLM call. It can use Plainsay Cloud, your own key, a compatible local endpoint, or be switched off. By default it preserves the speaker's language. With an active Cloud subscription and a style-aware BYOK or compatible local cleaner selected, the same pass can instead Auto-translate each new dictation to a target chosen from the user's spoken languages or English. Email mode is off by default and can lay greetings, paragraphs, and sign-offs out as an email when the target is a supported mail app or recognized webmail compose window. Translation and email layout are independent, so an email can be dictated in one language and written in another.

Until v0.2.33 there was a provider gap here: only BYOK and compatible local cleaners received the style prompt, so Auto-translate and email layout did nothing on the built-in Cloud route — the very route their subscription pays for. The Cloud request now carries the layout and translation target as structured fields, and the server composes the prompt wording from them; the client never supplies prompt text for a key Plainsay pays for.

The design constraint that actually matters: a failed Polishing call falls back to the raw transcript. If Polishing is off, disabled, lacks a configured credential, or uses a cleaner that does not receive the selected style, translation and email layout do not run. A configured provider has a bounded timeout; network, timeout, HTTP, and provider-output-limit failures fall back to raw text with a HUD notice. The prompt also forbids inventing an ending when a recording stops mid-sentence. TextCleaning is a protocol with one method — the whole pipeline doesn't know or care whether the concrete implementation is Gemini, OpenAI-compatible, Anthropic, or a no-op passthrough when Polishing is disabled.

Getting text into someone else's app

Each recording is bounded at 10 minutes. When it reaches that boundary, Plainsay stops capture automatically, keeps the full bounded recording, shows a notice, and continues through transcription and insertion.

Audio diagnostics compare how long recording was active with the duration represented by the captured samples, and flag a meaningful shortfall or an interrupted audio device. Settings › About provides a copyable diagnostics command for investigating that path.

Text lands via the pasteboard and a synthetic ⌘V, not Apple's Accessibility text-insertion API. This works across many native apps, Electron apps, web views, and terminals. If Plainsay cannot identify a safe destination, the text remains on the clipboard so it can be pasted manually.

Input Monitoring is used to notice the configured dictation shortcut, the global translation toggle, and Escape while recording. Plainsay does not store typed text or include keystrokes in a transcript. On the first request it brings the macOS permission prompt to the front. If permission is requested again after macOS no longer has a prompt to show, Plainsay opens the relevant System Settings pane. If macOS quits the app while applying Accessibility or Input Monitoring, Setup resumes at the same step after relaunch.

The cost is briefly borrowing the real clipboard: Plainsay snapshots the existing contents before the write and attempts to restore them after insertion. If it cannot identify a safe destination, it leaves the transcript on the clipboard for manual paste.

Read the rest

Why on-device — the privacy case, stated precisely. Benchmark — WER and latency for both engines, reproducible with the documented build-and-run commands. Plainsay client source — MIT licensed; third-party licenses are listed in THIRD_PARTY_NOTICES.md.

Read the code, or just use it.

Download for macOS Free · MIT licensed · macOS 14+, Apple silicon