MumbleFlowGet MumbleFlow — $5

Voice typing / macOS

Speech to text for Mac that turns talking into finished writing

MumbleFlow is a native, local-first dictation app for people who think faster out loud. Activate it where you write, speak naturally, and receive a stable result with optional corrections, punctuation, paragraphs, and list formatting.

MumbleFlow currently requires an Apple Silicon Mac, macOS 26 or later, and English speech. Recognition quality still depends on the microphone, environment, accent, vocabulary, and what was said.

Apple Silicon · macOS 26+ · English · Models download during setup

Dictation History

We’re moving the launch to Thursday afternoon to give onboarding one more review. Thanks for the extra care as we get this ready.

2 minutes ago

1. Review the onboarding flow.
2. Update the launch checklist.
3. Share the final note with the team.

18 minutes ago

Can you help me think through the simplest version of this idea? I want the first step to feel obvious.

1 hour ago

Hold fn to dictate. Double-tap to keep talking.
The words you wanted, ready where you work.App interface preview · Sample content

From voice to cursor

A speech-to-text workflow that stays out of the way

Mac voice typing should not require opening a recorder, exporting an audio file, or moving text between tools. MumbleFlow lives in the menu bar and gives you two Fn gestures for short and long-form dictation. A compact frosted pill confirms that capture is active, its waveform responds to microphone level, and a processing indicator replaces the waveform after you stop.

  1. 01

    Focus your destination

    Place the cursor in the message, document, browser editor, or other editable field where the finished writing should go. MumbleFlow targets the field that is focused at insertion time.

  2. 02

    Speak in your own words

    Hold Fn for a quick thought or double-tap Fn to record hands-free. Ramble, pause, introduce a list, or correct yourself aloud instead of mentally composing every sentence first.

  3. 03

    Finish or cancel

    Release Fn to finish a held recording, press Fn once to finish a locked recording, or press Escape to cancel. MumbleFlow processes the capture only after you finish.

  4. 04

    Use one stable result

    The completed text is inserted at the cursor, copied to the clipboard, and optionally retained in encrypted Dictation History. Unstable partial words are not typed into the destination as you speak.

Quiet sound cues mark the start, locked-recording, and finish states. If a keyboard reserves Fn, the fallback shortcut provides the same capture flow. The offline dictation setup guide walks through the microphone, Accessibility, model, and shortcut checks.

Beyond verbatim dictation

Keep the meaning. Remove the work of cleaning it up.

Built-in dictation often treats every spoken fragment as part of the final document. MumbleFlow separates speech recognition from writing polish. The recognizer first produces a transcript; an optional on-device language model then edits that transcript conservatively.

01REVISE

Spoken corrections

A phrase such as “five—actually six” can be reduced to the corrected thought rather than preserving both versions. The goal is a faithful edit, not a new idea.

02STRUCTURE

Punctuation and paragraphs

Polishing adds sentence boundaries, capitalization, questions, and paragraph breaks so a natural monologue reads like intentional writing.

03FORMAT

Numbered tasks

When you clearly introduce a sequence of tasks, MumbleFlow can turn the spoken enumeration into a readable numbered list instead of one dense paragraph.

04CONTEXT

Personal vocabulary

A local personal dictionary gives the recognizer context for the names, acronyms, and specialist terms you actually use without sending that vocabulary to a hosted speech API.

05CONTROL

Short-phrase fast path

Very short utterances can skip language-model rewriting when semantic cleanup would add little value, while deterministic text normalization still runs.

06OPTIONAL

Raw-text option

Turn polishing off when a close transcript matters more than a composed result. MumbleFlow then applies deterministic whitespace cleanup rather than a semantic rewrite.

Polishing is instructed to preserve meaning, tone, names, numbers, links, and substantive detail. It is useful for messages, emails, working notes, task lists, and first drafts—but it is not a guarantee that every correction or uncommon term will be interpreted perfectly. Review consequential names, dates, numbers, and instructions before sending.

System-wide dictation

Hold for one sentence. Lock for a longer thought.

Momentary

Hold Fn

Capture begins on key-down so the first word is not intentionally delayed. Keep Fn held while speaking, then release to stop immediately and begin recognition.

Hands-free

Double-tap Fn

Lock recording when the thought is longer than a sentence. Speak without holding the keyboard, then press Fn once to stop. Escape cancels either recording mode.

After processing, MumbleFlow first attempts direct selected-text replacement through macOS Accessibility. It can fall back to a standard paste event when a browser or cross-platform editor exposes focus differently. Because applications control their own editors, the result is also left on the clipboard for recovery. See dictation and insertion troubleshooting for practical checks.

Dictation History keeps recent results available when it is enabled. Unpinned entries expire after seven days by default; a useful entry can be pinned or copied, and unpinned entries can be cleared. For the complete data path, read how private voice to text works.

Local-first pipeline

Recognition and writing cleanup run on your Mac

MumbleFlow uses FluidAudio’s Silero voice-activity detector to find spoken sections and an English Parakeet TDT 0.6B v2 Core ML model for recognition. The model assets download during setup, are verified, and then run on the Apple Silicon Mac rather than calling a hosted transcription API for each utterance.

Longer text can be polished with Apple’s on-device Foundation Models when they are available. If Apple Intelligence is unavailable, an optional local Qwen3 model can provide semantic cleanup. Settings show which processing path is active, and polishing can be turned off without disabling raw transcription.

Temporary dictation audio is deleted after processing unless you enable opt-in diagnostics to retain the latest test recording. The transcript, clipboard, destination app, and optional history each have different privacy implications, so local recognition does not make text private after you insert it into a third-party service. The privacy and retention page explains those boundaries.

Before you download

Built for a specific Mac workflow—not every device

MumbleFlow’s current release is for Apple Silicon Macs running macOS 26 or later, with English recognition. It requires Microphone access for capture and Accessibility access for global Fn handling and text insertion. The native Swift, SwiftUI, and AppKit application is distributed outside the Mac App Store as a signed and notarized DMG.

Choose it when you want local-first voice typing, system-wide controls, optional writing polish, a personal dictionary, recoverable clipboard output, and encrypted local history. It is not currently a Windows, Intel Mac, mobile, multilingual, or cloud-sync product.

If your primary need is long-form calls rather than quick writing, explore Oat AI meeting notes and speaker-aware meeting transcription. For the whole product, visit the complete MumbleFlow feature list.

Questions

Frequently asked questions

How do I start speech to text on my Mac with MumbleFlow?

Hold Fn while you speak and release it to finish, or double-tap Fn to lock recording for a longer passage and press Fn once to stop. If Fn conflicts with your keyboard or macOS settings, MumbleFlow also provides a configurable fallback shortcut that defaults to Option-Space.

Does MumbleFlow type every filler word and false start?

Not when polishing is enabled. For longer dictation, MumbleFlow can remove fillers and abandoned starts, resolve explicit spoken corrections, add punctuation and paragraph breaks, and format clearly introduced lists. Polishing can also be disabled when you want a closer raw transcript.

Can MumbleFlow voice typing work without an internet connection?

Yes, after the app and English speech assets have been downloaded. Recognition runs locally with Parakeet TDT, and polishing can use Apple Foundation Models or an optional local Qwen3 model. Downloads, software updates, and purchases still require a connection; external source connectors are not enabled in the current download.

Can MumbleFlow insert speech-to-text results into browser editors?

MumbleFlow is designed to insert the finished result into the editable field focused when processing ends. It uses macOS Accessibility where possible and a clipboard-and-paste fallback for editors that expose text differently. The final result is also copied so it remains available if a particular site rejects automated insertion.

What happens to my dictation audio and text?

Temporary dictation audio is deleted after processing unless you enable opt-in diagnostics to retain the latest test recording. Dictated text can be kept in encrypted local Dictation History for seven days by default; you can disable history, pin an entry, copy it, or clear unpinned entries.

Which Macs and languages does MumbleFlow support?

The current MumbleFlow release supports English on Apple Silicon Macs running macOS 26 or later. Intel Macs, older macOS releases, Windows, and additional recognition languages are not supported in this release.

Keep exploring

The rest of MumbleFlow

Looking for technical detail? Read the local speech-to-text guide.