Guide
How accurate is AI transcription, and what affects it?
Modern AI transcription is good enough that clean, clearly-spoken audio comes back with only occasional errors, and it keeps getting better. It is not perfect, and accuracy drops with anything that makes speech harder to hear: background noise, a poor microphone, strong or unfamiliar accents, people talking over each other, and specialist jargon or names the model has not seen much.
The biggest lever you control is the audio itself. A decent microphone, a quiet room, and one person speaking at a time will do more for accuracy than any single feature. Recording each speaker on a separate track also helps, because the model is not trying to untangle overlapping voices.
Good, and better than it was
Modern AI transcription is at the point where clean, clearly-spoken audio comes back with only the odd error, and the models keep improving. A one-on-one call with good microphones, or an online meeting where everyone takes turns, will usually transcribe well enough to use with a quick read-through.
It is not perfect, and it never will be, because some of what hurts accuracy is in the audio rather than the model. Knowing what those things are tells you where a transcript is likely to slip, and what you can do about it.
What trips it up
A few things reliably raise the error rate.
- Poor audio. Background noise, a distant or low-quality microphone, and echo all make speech harder to make out, for a model just as for a person.
- Overlapping speech. When two people talk at once, the model has to guess at a tangle. This is the single most common source of garbled lines in meeting transcripts.
- Accents and unfamiliar voices. Models handle common accents well and struggle more with ones they have seen less of.
- Jargon, names and acronyms. Everyday words come back reliably; specialist terms, product names and people’s names are where errors cluster.
What actually improves accuracy
The biggest lever is the audio you feed in, not any setting. A decent microphone, a quiet room, and one person speaking at a time will do more for a transcript than switching models. If you are recording a call, recording each speaker on a separate track helps too, because the model is transcribing one clear voice at a time instead of untangling overlapping ones.
After that, a quick read-through to fix the names and jargon is usually all a good transcript needs. Those errors cluster in predictable places, so correcting them is fast once you know to look.
On-device versus the cloud
There used to be a real accuracy gap between transcribing on your own machine and sending audio to a cloud service. For everyday meetings, that gap has largely closed. Current on-device models are strong, so for most audio the choice between local and cloud is now about privacy rather than quality. If your machine can run a good model locally, you are not trading much away to keep the audio on it.
FAQ
What hurts transcription accuracy the most?
Poor audio, above everything: background noise, a distant or low-quality microphone, and people talking over each other. Accents, technical jargon and unusual names also raise the error rate.
How do I get a more accurate transcript?
Improve the input. Use a good microphone, record in a quiet space, and have one person speak at a time. Recording speakers on separate tracks helps the model keep them apart.
Is on-device transcription less accurate than the cloud?
Not meaningfully for most audio. On-device models like Whisper large-v3-turbo are current and strong; the gap that used to exist has largely closed for everyday meetings.
Will it get names and jargon right?
Common words come back reliably; rare names, acronyms and specialist terms are where errors cluster. A quick read-through to fix those is usually all a good transcript needs.
Try Trace
Trace records and transcribes on your Mac, on-device, with no bot and no cloud. One-time purchase, no subscription.
Related guides
- Transcribe a recording you have: Turn an existing audio or video file into text, without uploading it.
- Transcript to summary: Go from a raw transcript to a short summary with clear action items.
- Offline transcription for Mac: Everything stays on your machine, and it works with the network off.