Why dictation always spells your name wrong
It is the coffee-cup thing, but on your screen. You say your name clearly and dictation writes back a spelling that is not yours. It is not mishearing you, it is playing the odds. Here is why, and how to make it stop.

Image: Pexels
You know the barista who writes a friendly, confident, completely wrong spelling of your name on the cup? You said 'Kris' and got 'Chris'. You said 'Zoey' and got 'Zoe'. Voice-to-text does the exact same thing, except it happens every single day, inside your emails, your messages, and your notes.
The frustrating part is that it feels personal. You enunciate. You slow down. You practically spell it out. And the little cursor calmly types back the wrong name again, as if it knows better than you do. Here is what is actually going on, and the one setting that ends it.
The coffee-cup problem, now on your screen
This is the everyday humiliation of dictation: it hands you back a name that is close enough to sound right out loud but wrong on the page. People report the same pattern over and over. On Apple's own forums, users describe dictation turning 'Zoey' into 'Zoe' and 'Kruse' into the wrong word, every time, no matter how they say it.
- You say Zoey, it writes Zoe (a real, documented complaint).
- You say a last name like Kruse and it swaps in a common word instead.
- You say Kris, Sofia, or Yesica, and get the crowd-favorite spelling: Chris, Sophia, Jessica.
- Your brand, your product, your colleague's name: all quietly 'corrected' into something else.
None of these are you speaking badly. They are the tool picking the spelling it sees most often and assuming that is the one you meant.
Why your name keeps losing
Think of it as the barista writing their favorite spelling of your name. That is an analogy, not a description of the machine, but it captures the real behavior: when a name can be spelled more than one way, the tool leans toward the version it has seen most. It is playing the odds, and your spelling is not the odds-on favorite.
The honest technical version comes straight from OpenAI, whose Whisper model powers a lot of modern dictation. Their own guidance says the model 'often does not recognize uncommon words or acronyms' and will misspell proper nouns unless you give it the correct spelling first. Their example: 'DALL-E' and 'GPT-3' came out as 'DALI' and 'GDP 3'. Your unusual name is just another uncommon word it has to guess at.
This is also why the usual instincts do not help. The tool is not mishearing the sound of your voice, it is choosing the more probable spelling for that sound. So speaking louder or more slowly does not move the needle, because volume was never the problem. The giveaway is that spelling it out letter by letter (say 'Steven, S-T-E-P-H-E-N') often does work, which shows the issue is spelling choice, not hearing.
The tell
If saying your name normally fails but spelling it letter by letter succeeds, the tool heard you fine. It just picked the wrong spelling. That is a vocabulary problem, not a microphone problem.
Spanish names get it worse
If you live your day in Spanish, this goes from annoying to a running joke. The 'most common spelling wins' logic is trained on a mountain of English text, so Spanish names get flattened toward whatever looks familiar to it. Jose loses its accent. Yesica becomes 'Jessica'. Xiomara, Nayeli, and plenty of others come back mangled or Anglicized.
Tildes are the first casualty. A dictation tool that types 'Jose' instead of 'José', or 'Nunez' instead of 'Núñez', is not being careless, it simply never treated your accented spelling as the likely one. Multiply that by every email, every client, every message you send, and it is a small daily tax on your own name.
The band-aids (and where they stop)
There are real workarounds built into your devices. They help, but each one is a patch on a single crack, not a fix for the wall.
Apple Text Replacement
On iPhone or Mac, go to Settings, then General, then Keyboard, then Text Replacement, and map a phrase to the exact spelling you want. Good for a handful of fixed terms.
Add a phonetic name in Contacts
Adding a pronunciation or phonetic field to a contact helps Siri and dictation get that person's name right. Useful, but it only covers people in your Contacts.
Spell it out in the moment
In regular dictation you can spell letter by letter, like 'Steven, S-T-E-P-H-E-N'. The dedicated hands-free 'Spelling Mode' is part of Voice Control, not standard Dictation, so do not expect it inside the everyday dictation flow.
On Windows, be precise about which tool you mean. The quick Win+H voice typing overlay still has no easy user dictionary. But Voice Access, Microsoft's fuller accessibility dictation tool on Windows 11 22H2 and later (open it with Ctrl+Win+S), now lets you add your own words via 'Add to Vocabulary', after 'Spell that' or 'Correct that', or from Voice Access settings. That capability rolled out through 2025, so if you only ever used Win+H, you may not know it exists.
Why band-aids run out
Text Replacement and Contacts fixes live on one device and cover fixed cases. Spelling things out mid-sentence breaks your flow. None of them travel with you across every app and machine. That is the real gap.
The real fix: teach it once
The durable answer is a custom vocabulary: a short list of the exact spellings you want, taught one time, applied to your transcripts everywhere. Instead of correcting the same name after the fact, you tell the tool up front, 'this is how these words are spelled,' and it stops guessing. OpenAI's own guidance points the same direction: feed the correct spellings in and the misspellings go away.
One caveat if you have heard of the 'just paste your names into the prompt' trick: Whisper only considers the final 224 tokens of that prompt, so a raw pasted word list is capped and fragile. A managed vocabulary you teach once, and that gets applied every time, is far more useful than re-pasting a list you keep growing.
How to build your vocabulary list
You do not need hundreds of entries. You need the ten to twenty words you actually say every week.
List the words you say weekly
Your own name, your family, your closest colleagues, your brand and product names, and the jargon of your field. Start with what comes up most.
Write the exact spelling
Include tildes and capitalization exactly as you want them: José, not Jose. Núñez, not Nunez. The tool copies what you give it.
Test it
Dictate a normal sentence using each word and confirm it lands right. Fix any that slipped.
Add as you go
When a new name keeps getting mangled, add it once. The list quietly gets better and you stop re-correcting.
| Correcting by hand | Teaching a vocabulary once | |
|---|---|---|
| Effort | Every single time | Once, then done |
| Travels across apps | No | Yes |
| Handles tildes and caps | Only if you retype | Yes, as taught |
| Your name over time | Still wrong | Spelled your way |
Do it once with Golem
Plenty of tools now offer a custom dictionary, and you should compare fairly. Superwhisper has a custom-vocabulary system, and Spokenly ships a custom dictionary its docs describe as working at the speech-recognition level. Both are legitimate options if you want to teach your names once.
Golem builds this in too. You teach it the exact spelling of a word one time, tildes and capitals included, and it uses that spelling in your transcripts across any app: your browser, your editor, your email, your chat, the ChatGPT and Claude web apps. Teach it 'Yesica', 'Xiomara', or your brand name once, and you stop fighting your own name. To be straight about how it works: Golem does not rewrite the speech model, it applies your spellings to the transcript it produces, so you get the outcome you wanted without a fragile paste-a-list ritual.
Frequently asked questions
Why does dictation always spell my name wrong?
Because when a name can be spelled more than one way, the tool leans toward the spelling it has seen most, not necessarily yours. OpenAI's Whisper guidance says the model often misspells uncommon words and proper nouns unless you give it the correct spelling first.
Does speaking louder or slower help?
Not really. The tool is not mishearing the sound, it is choosing the more probable spelling for that sound. The proof is that spelling it letter by letter often works, which means it heard you fine and just picked the wrong spelling.
Can I add my name to Windows dictation?
The quick Win+H voice typing overlay still has no easy user dictionary. But Windows 11 Voice Access (Ctrl+Win+S) now lets you add words via 'Add to Vocabulary', a capability that rolled out through 2025. They are two different tools.
How does Golem fix it?
You teach Golem the exact spelling of a word once, tildes and capitalization included, and it applies that spelling to your transcripts across any app. No retyping, no pasting a list every time.

