Home/Docs/Dictation basics

Dictation basics

One key, held. Everything else the app does is in service of that gesture, and the floating HUD tells you which stage you are in.

The loop

  1. Click where the text should go — any app, any text field.
  2. Hold the key (fn by default) and speak normally.
  3. Release. The app transcribes on your Mac and inserts the text at your cursor.

Because the microphone is live only while the key is down, there is no timeout to race and no mode to remember exiting. Pause mid-sentence to think for as long as you like; the app is still listening as long as you are still holding.

Reading the HUD

A small floating pill appears while you dictate. Its states, in order:

StateMeaning
Waveform movingRecording. The wave tracks your voice, so a flat line means the app is not hearing you.
Words appearingLive partial transcript — proof it is hearing you, refreshed about three times a second.
Spool spinningTranscribing on your Mac after you released the key.
CleaningOnly if AI cleanup is on. Skipped entirely when it is off.
ResultThe tail of what was inserted, shown briefly before the HUD hides itself.

Errors surface here too, each with a one-tap action: No microphone (opens Sound settings), Took too long — audio kept (retry), and Copied instead — press ⌘V (see How text is inserted).

Length limits

  • Minimum: about half a second. Anything shorter is treated as a mis-tap and discarded — the HUD says so rather than pasting noise.
  • Maximum: five minutes per take by default, adjustable to 1, 2, 5, or 10 minutes in Settings. When the cap is reached the app finishes the take and inserts it — your audio is never thrown away. The cap exists so a key that gets stuck down can't record forever.

The HUD counts down toward the cap while you hold, so a long take never ends as a surprise.

Pausing

The menu-bar menu has Pause for 1 hour, which stops the app from responding to the hotkey — useful when you are handing your Mac to someone, or when the key you chose is needed for something else. The menu-bar icon shows the paused state, and you can resume from the same menu at any time.

Speak in phrases, not words

Whisper uses surrounding context to choose between homophones, so dictating a whole clause at natural speed is markedly more accurate than word-by-word delivery. More in Improving accuracy.

Punctuation

Whisper adds punctuation from sentence structure — you generally do not need to do anything. Note that OpenVoiceFlow 0.5.5 has no spoken-punctuation commands: saying “comma” produces the word “comma”, not the mark. If you want filler words removed and grammar tidied, that is what AI cleanup is for.

Documentation for OpenVoiceFlow 0.5.5 · last updated 2026-07-31. Something wrong or missing here? That is a bug — tell us at shimoverse@gmail.com.