Voice typing / macOS
Speech to text for Mac that turns talking into finished writing
MumbleFlow is a native, local-first dictation app for people who think faster out loud. Activate it where you write, speak naturally, and receive a stable result with optional corrections, punctuation, paragraphs, and list formatting.
Apple Silicon · macOS 26+ · English · Models download during setup
Product sync
Let’s make the first run feel a little simpler. What still needs to happen before we launch?
We’re moving the launch to Thursday afternoon so we can give onboarding one more review.
I’ll take the onboarding review. I can have the updated checklist ready by Wednesday.
Great. Let’s keep the launch note short and calm. Mention the extra time for onboarding.
One thing is still open: do we send the update to everyone, or start with the pilot group?
Topics
Launch timing
A little more time for a careful first impression.
Onboarding review
A final pass through the first-run experience.
Timeline
The team reviewed the remaining launch work and agreed to make room for one more onboarding check.
Decisions
Move the launch to Thursday afternoon.
Next steps
Maya will finish the onboarding checklist by Wednesday.
Open questions
Should the update go to everyone or the pilot group first?
What you missed
Last 5 minutesA calm launch update
Keep the note short and mention the extra time for onboarding.
Next steps
Maya will finish the checklist by Wednesday.
Dictation History
We’re moving the launch to Thursday afternoon to give onboarding one more review. Thanks for the extra care as we get this ready.
2 minutes ago1. Review the onboarding flow.
2. Update the launch checklist.
3. Share the final note with the team.
Can you help me think through the simplest version of this idea? I want the first step to feel obvious.
1 hour agoSearch everything
The launch is moving to Thursday afternoon, giving the team time for one more onboarding review.
From voice to cursor
A speech-to-text workflow that stays out of the way
Mac voice typing should not require opening a recorder, exporting an audio file, or moving text between tools. MumbleFlow lives in the menu bar and gives you two Fn gestures for short and long-form dictation. A compact frosted pill confirms that capture is active, its waveform responds to microphone level, and a processing indicator replaces the waveform after you stop.
- 01
Focus your destination
Place the cursor in the message, document, browser editor, or other editable field where the finished writing should go. MumbleFlow targets the field that is focused at insertion time.
- 02
Speak in your own words
Hold Fn for a quick thought or double-tap Fn to record hands-free. Ramble, pause, introduce a list, or correct yourself aloud instead of mentally composing every sentence first.
- 03
Finish or cancel
Release Fn to finish a held recording, press Fn once to finish a locked recording, or press Escape to cancel. MumbleFlow processes the capture only after you finish.
- 04
Use one stable result
The completed text is inserted at the cursor, copied to the clipboard, and optionally retained in encrypted Dictation History. Unstable partial words are not typed into the destination as you speak.
Quiet sound cues mark the start, locked-recording, and finish states. If a keyboard reserves Fn, the fallback shortcut provides the same capture flow. The offline dictation setup guide walks through the microphone, Accessibility, model, and shortcut checks.
Beyond verbatim dictation
Keep the meaning. Remove the work of cleaning it up.
Built-in dictation often treats every spoken fragment as part of the final document. MumbleFlow separates speech recognition from writing polish. The recognizer first produces a transcript; an optional on-device language model then edits that transcript conservatively.
Spoken corrections
A phrase such as “five—actually six” can be reduced to the corrected thought rather than preserving both versions. The goal is a faithful edit, not a new idea.
Punctuation and paragraphs
Polishing adds sentence boundaries, capitalization, questions, and paragraph breaks so a natural monologue reads like intentional writing.
Numbered tasks
When you clearly introduce a sequence of tasks, MumbleFlow can turn the spoken enumeration into a readable numbered list instead of one dense paragraph.
Personal vocabulary
A local personal dictionary gives the recognizer context for the names, acronyms, and specialist terms you actually use without sending that vocabulary to a hosted speech API.
Short-phrase fast path
Very short utterances can skip language-model rewriting when semantic cleanup would add little value, while deterministic text normalization still runs.
Raw-text option
Turn polishing off when a close transcript matters more than a composed result. MumbleFlow then applies deterministic whitespace cleanup rather than a semantic rewrite.
Polishing is instructed to preserve meaning, tone, names, numbers, links, and substantive detail. It is useful for messages, emails, working notes, task lists, and first drafts—but it is not a guarantee that every correction or uncommon term will be interpreted perfectly. Review consequential names, dates, numbers, and instructions before sending.
System-wide dictation
Hold for one sentence. Lock for a longer thought.
Momentary
Hold Fn
Capture begins on key-down so the first word is not intentionally delayed. Keep Fn held while speaking, then release to stop immediately and begin recognition.
Hands-free
Double-tap Fn
Lock recording when the thought is longer than a sentence. Speak without holding the keyboard, then press Fn once to stop. Escape cancels either recording mode.
After processing, MumbleFlow first attempts direct selected-text replacement through macOS Accessibility. It can fall back to a standard paste event when a browser or cross-platform editor exposes focus differently. Because applications control their own editors, the result is also left on the clipboard for recovery. See dictation and insertion troubleshooting for practical checks.
Dictation History keeps recent results available when it is enabled. Unpinned entries expire after seven days by default; a useful entry can be pinned or copied, and unpinned entries can be cleared. For the complete data path, read how private voice to text works.
Local-first pipeline
Recognition and writing cleanup run on your Mac
MumbleFlow uses FluidAudio’s Silero voice-activity detector to find spoken sections and an English Parakeet TDT 0.6B v2 Core ML model for recognition. The model assets download during setup, are verified, and then run on the Apple Silicon Mac rather than calling a hosted transcription API for each utterance.
Longer text can be polished with Apple’s on-device Foundation Models when they are available. If Apple Intelligence is unavailable, an optional local Qwen3 model can provide semantic cleanup. Settings show which processing path is active, and polishing can be turned off without disabling raw transcription.
Temporary dictation audio is deleted after processing unless you enable opt-in diagnostics to retain the latest test recording. The transcript, clipboard, destination app, and optional history each have different privacy implications, so local recognition does not make text private after you insert it into a third-party service. The privacy and retention page explains those boundaries.
Before you download
Built for a specific Mac workflow—not every device
MumbleFlow’s current release is for Apple Silicon Macs running macOS 26 or later, with English recognition. It requires Microphone access for capture and Accessibility access for global Fn handling and text insertion. The native Swift, SwiftUI, and AppKit application is distributed outside the Mac App Store as a signed and notarized DMG.
Choose it when you want local-first voice typing, system-wide controls, optional writing polish, a personal dictionary, recoverable clipboard output, and encrypted local history. It is not currently a Windows, Intel Mac, mobile, multilingual, or cloud-sync product.
If your primary need is long-form calls rather than quick writing, explore Oat AI meeting notes and speaker-aware meeting transcription. For the whole product, visit the complete MumbleFlow feature list.
Questions
Frequently asked questions
How do I start speech to text on my Mac with MumbleFlow?
Hold Fn while you speak and release it to finish, or double-tap Fn to lock recording for a longer passage and press Fn once to stop. If Fn conflicts with your keyboard or macOS settings, MumbleFlow also provides a configurable fallback shortcut that defaults to Option-Space.
Does MumbleFlow type every filler word and false start?
Not when polishing is enabled. For longer dictation, MumbleFlow can remove fillers and abandoned starts, resolve explicit spoken corrections, add punctuation and paragraph breaks, and format clearly introduced lists. Polishing can also be disabled when you want a closer raw transcript.
Can MumbleFlow voice typing work without an internet connection?
Yes, after the app and English speech assets have been downloaded. Recognition runs locally with Parakeet TDT, and polishing can use Apple Foundation Models or an optional local Qwen3 model. Downloads, software updates, and purchases still require a connection; external source connectors are not enabled in the current download.
Can MumbleFlow insert speech-to-text results into browser editors?
MumbleFlow is designed to insert the finished result into the editable field focused when processing ends. It uses macOS Accessibility where possible and a clipboard-and-paste fallback for editors that expose text differently. The final result is also copied so it remains available if a particular site rejects automated insertion.
What happens to my dictation audio and text?
Temporary dictation audio is deleted after processing unless you enable opt-in diagnostics to retain the latest test recording. Dictated text can be kept in encrypted local Dictation History for seven days by default; you can disable history, pin an entry, copy it, or clear unpinned entries.
Which Macs and languages does MumbleFlow support?
The current MumbleFlow release supports English on Apple Silicon Macs running macOS 26 or later. Intel Macs, older macOS releases, Windows, and additional recognition languages are not supported in this release.
Keep exploring
The rest of MumbleFlow
All features
A complete tour of dictation, Oat meeting notes, search, personalization, privacy, and controls.
Oat AI meeting notes
Reviewable, actionable meeting notes with live catch-up, speaker memory, timelines, and source citations.
Meeting transcription
Separate microphone and computer audio, readable speaker turns, timestamp playback, and grounded notes.
Local AI search
Private hybrid search and grounded answers that take you directly to the supporting source.
Privacy and security
A plain-language account of local processing, encryption, retention, network access, and deletion.
About MumbleFlow
Why MumbleFlow exists, how it is built, and the product principles behind it.
Looking for technical detail? Read the local speech-to-text guide.