Guides
Fixing dictation mistakes: it is usually not the model
Before changing app, change four things. Almost every accuracy complaint traces to the microphone, the room, how you are speaking, or words the app has never been taught.
Last updated
1. The microphone, which is most of it
Speech recognition quality is dominated by the signal it receives. A model cannot recover information that never reached it, and the built-in microphone on a laptop is eighteen inches away, pointed at the ceiling, next to a fan.
This is the single biggest lever and the one people try last, because changing an app feels like a solution and moving a microphone does not.
| Setup | Typical result |
|---|---|
| Built-in MacBook microphone, lid open, normal sitting distance | Workable in a quiet room, poor in any other |
| AirPods or similar | Better — much closer to your mouth. The usual first upgrade |
| A wired headset with a boom microphone | Substantially better, and lets you speak quietly |
| A desk microphone 15–20cm away, off-axis | Excellent, if the room is reasonable |
| Any microphone in a room with hard walls and no soft furnishings | Echo defeats all of the above |
2. The room
Echo is worse than noise. A hard-surfaced room sends the model your voice twice, slightly apart, and it has to reconcile them. Bare kitchens and bathrooms are the classic offenders; so are minimalist home offices with a desk, a window and nothing soft.
- Anything soft helps. A rug, curtains, a bookshelf, a coat on a hook. You are not building a studio, you are absorbing reflections.
- Move away from the hard wall, not towards it.
- Constant noise is easier than intermittent. Fans and traffic are handled better than a conversation happening nearby, because speech competes with speech.
- A close microphone beats a treated room, so if you can only fix one, fix the microphone.
3. How you are speaking
Three habits cause most of the remaining errors, and all three are the opposite of what people instinctively do when dictation goes wrong.
- Do not slow down The instinct after an error is to speak more slowly and deliberately. This makes it worse. Models are trained on natural speech, and exaggerated word-by-word delivery is unlike anything in the training data.
- Speak complete phrases Recognition uses surrounding words to decide what a word was. "I'll be there at eight" is easier to get right than "eight" alone, because the context disambiguates it.
- Do not trail off Volume dropping at the end of a sentence is very common and costs the last three words. Keep the level up to the full stop.
- Pause before you start Half a second between pressing the shortcut and speaking. Starting instantly clips the first word, and the first word sets the context for everything after it.
4. Words it has never been taught
Once audio and delivery are sorted, what remains is almost entirely proper nouns and jargon. These are not accuracy failures — the model heard you correctly and does not know the word.
Fifteen vocabulary entries fixes it, permanently. Full method at building a dictation vocabulary.
Specific symptoms
| Symptom | Most likely cause | Fix |
|---|---|---|
| The first word is missing | You started speaking as you pressed | Half a second of pause first |
| The last few words are missing or wrong | Trailing off, or releasing too early | Keep the volume up; release a beat after you finish |
| Numbers come out as words, or vice versa | A formatting preference, not an error | Check the app's number handling setting |
| Homophones are wrong (their/there) | The model chose by context and got it wrong | Read-back. Nothing else fixes this |
| Names are consistently wrong | Not taught | Custom vocabulary |
| Random unrelated words appear | Background speech being picked up | Close microphone, or a quieter room |
| Everything is wrong | Wrong input device selected | System Settings → Sound → Input |
| It is accurate but feels slow | Dictating word by word | Speak whole thoughts |
The read-back is not optional
No dictation setup reaches perfect, and the errors that survive are the dangerous ones — plausible words in grammatical sentences that mean something other than what you said. A spell checker will not catch "can" where you said "can't".
Ten seconds of reading before sending is part of the workflow. Dictating three sentences takes eight seconds; typing them takes forty-five. The arithmetic survives the read-back comfortably.
Questions
Why is Mac dictation so inaccurate for me?
In order of likelihood: the microphone is too far away, the room echoes, you are speaking word by word rather than in phrases, or the words it gets wrong are names it has never been taught. The model itself is usually the last thing to suspect.
Does a better microphone really help that much?
More than anything else on this page. Recognition works on the signal it receives, and a laptop microphone eighteen inches away in a room with a fan is a poor signal. Even inexpensive earbuds with a microphone closer to your mouth are a large improvement.
Should I speak slowly for better accuracy?
No — this is the most common wrong instinct. Models are trained on natural speech. Slow, exaggerated, word-by-word delivery is unlike the training data and produces worse results. Speak normally, in complete phrases.
Can I correct a dictation by voice?
Not in most modern dictation apps, which expect keyboard correction. macOS Voice Control is the tool designed for voice-driven editing on a Mac if that is essential for you.
How do I find out what I actually said?
An app with a local history that keeps the audio lets you play it back. Usually it turns out you said something slightly different from what you remember, which is more useful to know than any setting.