Oat / meeting transcript
Meeting transcription with speaker context and evidence you can open
Oat captures your microphone and selected computer audio without retaining video, then turns the meeting into a timestamped, speaker-aware transcript. Its source IDs connect later notes, decisions, and action items to the words behind them.
Apple Silicon · macOS 26+ · English · Models download during setup
Product sync
Let’s make the first run feel a little simpler. What still needs to happen before we launch?
We’re moving the launch to Thursday afternoon so we can give onboarding one more review.
I’ll take the onboarding review. I can have the updated checklist ready by Wednesday.
Great. Let’s keep the launch note short and calm. Mention the extra time for onboarding.
One thing is still open: do we send the update to everyone, or start with the pilot group?
Topics
Launch timing
A little more time for a careful first impression.
Onboarding review
A final pass through the first-run experience.
Timeline
The team reviewed the remaining launch work and agreed to make room for one more onboarding check.
Decisions
Move the launch to Thursday afternoon.
Next steps
Maya will finish the onboarding checklist by Wednesday.
Open questions
Should the update go to everyone or the pilot group first?
What you missed
Last 5 minutesA calm launch update
Keep the note short and mention the extra time for onboarding.
Next steps
Maya will finish the checklist by Wednesday.
Dictation History
We’re moving the launch to Thursday afternoon to give onboarding one more review. Thanks for the extra care as we get this ready.
2 minutes ago1. Review the onboarding flow.
2. Update the launch checklist.
3. Share the final note with the team.
Can you help me think through the simplest version of this idea? I want the first step to feel obvious.
1 hour agoSearch everything
The launch is moving to Thursday afternoon, giving the team time for one more onboarding review.
Explicit recording
Listen to the meeting on your Mac—without a bot or video capture
Oat does not need to join a call as a participant. After you press Record Meeting, ScreenCaptureKit captures the audio produced by the selected meeting application or display, while the microphone records your side as a separate stream. Video is neither captured nor retained.
- 01
Start deliberately
Choose Record Meeting, give the recording a title, and select the computer-audio source. Oat never begins because a calendar event started or a call app opened.
- 02
Keep channels separate
Your microphone and the selected computer audio are captured as separate timestamped streams. The microphone establishes which speech belongs to the local user.
- 03
Transcribe speech chunks
Voice activity detection closes completed speech chunks and sends them through the local recognizer while the meeting continues, building an ordered transcript over time.
- 04
Review and name
After the recording, correct uncertain speaker labels and review important names, numbers, decisions, and tasks against the cited transcript and retained audio.
System audio capture requires Screen & System Audio Recording permission; microphone capture requires Microphone permission. Calendar information may suggest a meeting title or participant names when a future connector is configured, but it does not start a recording. The current download does not enable Google or Slack connectors.
During the call
A readable transcript grows as completed speech segments arrive
FluidAudio’s voice-activity detection identifies completed speech chunks rather than waiting for the entire meeting to end. The local English Parakeet speech recognizer transcribes those chunks, and Oat orders them using the stream timestamps. Processing time can vary with meeting length, speech patterns, audio quality, and Mac load.
Separate local and remote audio
Because the microphone is not mixed into computer audio at capture time, Oat can label your turns as You before it analyzes who spoke remotely.
Speaker-aware turns
FluidAudio diarization divides mixed remote audio into speaker segments, which are aligned with recognition timestamps and merged with the microphone transcript.
Correctable names
A remembered name is suggested only for a sufficiently similar local voice profile. An uncertain match remains Speaker 1 until you label it rather than presenting a guess as fact.
Timestamped context
Every transcript segment carries its time range, channel, speaker or profile, text, and a source ID that downstream notes can cite.
Playback at the source
Select a citation to return to the relevant transcript timestamp and hear the retained encrypted recording while that audio is still available.
Searchable transcript
Finalized meeting segments feed MumbleFlow’s local meeting index so you can find a passage or ask a grounded question later.
Recording continues while the transcript is assembled. If you lose the thread, What did I miss? uses the preceding five minutes of finalized transcript to organize the recent topics, decisions, questions, and action items without stopping capture.
Who said what
Your microphone is You. Remote voices remain correctable.
Separating microphone and computer-audio channels solves one important attribution problem: Oat knows which speech came from the local microphone. For mixed remote audio, diarization estimates where one speaker ends and another begins. It aligns those regions with recognized words before merging them into the meeting timeline.
Oat keeps confirmed voice profiles locally. When a later speaker embedding is similar enough to a remembered profile, it can suggest that person’s name. If the match is below a conservative threshold, the transcript keeps a neutral label such as Speaker 1 and asks for confirmation rather than attaching a confident but unsupported identity.
You can correct a label and remember that mapping for future meetings. Similar voices, overlapping speech, poor audio, or a different microphone can still produce mistakes, so attribution should be checked before assigning an owner to a decision or task. For the broader notes workflow, see Oat AI meeting notes.
Grounded meeting notes
A generated claim should lead back to the transcript
A transcript segment is more than displayed text. It includes a meeting ID, start and end times, audio channel, speaker or voice-profile reference, raw and corrected text, and a stable source ID. Those IDs let MumbleFlow attach evidence to downstream notes.
Timeline events
Open the segment that establishes what happened and when.
Decisions
Return to the discussion supporting the recorded decision.
Action items
Check the language supporting the task, owner, or due date; unsupported owners and dates should not be added.
Open questions
See the original question or unresolved exchange in context.
Search answers
Follow cited results back to the matching meeting passage.
Selecting a citation opens the meeting at its timestamp and can play the corresponding local audio while the recording is retained. If the source does not support an answer, MumbleFlow’s grounded search is designed to say it could not find the information rather than fabricate an uncited response. Learn how retrieval works on the local AI search page.
Local meeting record
Encrypted audio expires; the transcript stays useful
Meeting recordings are encrypted on the Mac using a key protected through the macOS Keychain. Audio expires after 30 days by default, and Settings exposes storage use plus immediate deletion controls. The transcript and cited notes can remain after audio deletion, but their citations can no longer play a recording that no longer exists.
Speaker profiles, transcript text, meeting notes, and the local search index also remain on the device until their applicable deletion controls are used. MumbleFlow has no product account, content server, cloud sync, ambient recording, or automatic meeting start in this release.
Read the privacy, encryption, and retention details, compare all meeting and dictation capabilities on the features page, or use the support guide to troubleshoot system-audio routing.
Questions
Frequently asked questions
Can MumbleFlow transcribe computer audio from a meeting?
Yes. When you explicitly start an Oat meeting, MumbleFlow uses ScreenCaptureKit to capture selected Mac system or meeting-app audio and records your microphone as a separate timestamped stream. It does not retain video.
Does MumbleFlow automatically record every meeting?
No. Oat has no ambient buffer, meeting bot, or automatic calendar start. Recording begins only after you choose Record Meeting and ends when you stop it.
How does Oat know which speaker is me?
The microphone channel is treated as the local user. Oat diarizes the mixed remote-audio channel into speaker turns, can suggest a remembered name when a local voice-profile match clears a conservative threshold, and otherwise uses a generic label until you confirm the person.
Can I correct a speaker label?
Yes. You can relabel a speaker after the meeting, and MumbleFlow stores the confirmed local speaker profile so it can make a more useful suggestion in later recordings. You should still review speaker attribution when the distinction matters.
What is attached to a meeting-transcript citation?
Generated timeline events, decisions, action items, and other meeting statements retain one or more transcript-segment source IDs. Selecting a citation opens the meeting at the supporting timestamp and can play the retained local audio.
How long is meeting audio retained?
Meeting recordings are encrypted locally and expire after 30 days by default. Transcripts and notes can remain after the audio expires, and storage controls let you delete recordings immediately.
Keep exploring
The rest of MumbleFlow
All features
A complete tour of dictation, Oat meeting notes, search, personalization, privacy, and controls.
AI dictation for Mac
Natural voice typing that removes fillers, catches corrections, formats lists, and inserts finished text.
Oat AI meeting notes
Reviewable, actionable meeting notes with live catch-up, speaker memory, timelines, and source citations.
Local AI search
Private hybrid search and grounded answers that take you directly to the supporting source.
Privacy and security
A plain-language account of local processing, encryption, retention, network access, and deletion.
About MumbleFlow
Why MumbleFlow exists, how it is built, and the product principles behind it.
Looking for technical detail? Read the local speech-to-text guide.