If you have been comparing dictation software, you have probably collected a handful of percentages — 99%, 97%, 95% — and found that they do not help you choose. That is not because the numbers are dishonest. It is because a single accuracy figure answers a different question from the one you are asking.
You want to know: how well will this understand me, at my desk, with my microphone? A published accuracy figure answers: how well did this engine transcribe one particular set of recordings that neither of us has heard? Those two questions can have very different answers.
This page explains how the measurement works, so you can read any vendor's number — including any you see about us — and know what it is worth.
Word error rate, in plain English
Speech recognition accuracy is normally reported as word error rate, usually written WER. It is the share of words that came out wrong. Lower is better, and "95% accurate" is just another way of writing "5% WER".
What counts as wrong is more specific than most people expect. There are three kinds of error, and they are added together:
- Substitutions — a word came out as a different word. You said "their", it typed "there".
- Deletions — a word you said is missing from the transcript.
- Insertions — a word appears that you never said. Background speech and breath noise cause these.
Add the three together and divide by the number of words you actually spoke. Say you dictate a 100-word paragraph and the transcript has four wrong words, one missing word and one extra word: that is six errors in 100 words, a WER of 6%, or "94% accurate".
One consequence is worth knowing, because it explains a lot of confusing marketing: WER can exceed 100%. If the software inserts more words than you spoke — which happens in a noisy room — the error count can be larger than the word count. A figure that cannot go above 100 is not a word error rate.
Why the same software gives two people different results
The engine is only one of the things that decides your accuracy, and often not the most important one. The rest is your situation:
- Your microphone, and where it sits. This is usually the single biggest factor, and the cheapest to fix. A headset microphone a few centimetres from your mouth hears something quite different from a laptop microphone across the desk.
- The room. A television in the background, a fan, an open window, a hard-surfaced room with echo. All of it reaches the recogniser along with your voice.
- What you are dictating. Everyday prose is the easy case. Names, medical or legal terminology, part numbers, addresses and acronyms are much harder, because the software is choosing between words that sound alike.
- How you speak. Accent, pace, volume, whether you run sentences together, whether you trail off. None of this is a fault — it is simply variation the engine either handles or does not.
- Which model you run. Most on-device software ships several models of different sizes. A small model that runs comfortably on an older laptop will generally make more mistakes than a large one on a fast machine.
Two people running identical software, with identical settings, can land several percentage points apart on all of this alone. That is the gap no published figure can close for you.
What a benchmark number is actually measuring
Published word error rates come from standard research datasets — fixed collections of recordings with verified transcripts, so that different systems can be scored against the same material. That is genuinely useful for the people building the software: it is how you tell whether this month's model is better than last month's.
It is much less useful as a shopping guide, for three reasons:
- The recordings are not you. Benchmark audio has its own accents, its own microphones and its own subject matter, none of which are yours.
- Different vendors report different tests. Two figures measured on two datasets are not comparable, even when both are honestly reported. Read-aloud audiobook speech and spontaneous conversation produce very different scores from the same engine.
- The conditions are often left out. A number quoted without the dataset, the microphone and the speaking style is not a claim you can check.
When you do see a figure, the useful question is not "is it high?" but "measured on what?" If that is not stated, the figure is decoration.
Two different design philosophies
There is a real technical difference between the older and newer approaches to recognition, and it is worth understanding because it changes what "accuracy" even means for a given product.
Trained profiles. The traditional desktop approach builds a profile of your individual voice and vocabulary. You spend time up front, and you can keep teaching it — adding your own terminology, your colleagues' names, the words your profession uses. Dragon is the best-known product built this way, and Nuance's own materials claim up to 99% recognition accuracy. The trade-off is the setup time, and the fact that the profile is yours to maintain.
General models. The newer approach trains one large model on a very wide range of speech, once, and ships the same model to everyone. OpenAI's Whisper models are the best-known open example; the original Whisper paper describes training on around 680,000 hours of audio. There is no enrolment step — you install it and start talking — and equally there is no way to teach it your voice. What you get on the first day is broadly what you get later, aside from whatever the software adds around the model, such as a personal word list.
Neither approach is simply better. If you dictate specialist vocabulary all day, being able to train the software is worth real money. If you want to install something and use it this afternoon, an enrolment session is a cost with no upside. Which describes you is not something a percentage can tell you.
Test it on your own voice in ten minutes
This is the only benchmark that answers your question, and it is easy to run on any dictation tool with a trial — including the free ones already on your computer.
- Write out a short passage first — 150 to 200 words. Use your own writing, not a sample paragraph: your email style, your subject matter, the names you actually type.
- Read it aloud into the software the way you would normally speak, sitting where you normally sit.
- Compare the two side by side and count the substitutions, deletions and insertions. Divide by your word count.
- Repeat it once with a headset microphone if you did the first run on a built-in one. This single change is often worth more than switching products.
- Then do a realistic run. Dictate something unscripted — a real email — because speaking from a page is easier than speaking from your head, and the second number is the one you will live with.
Run the same passage through two or three tools and you will have something no review can give you: a comparison on your voice, your microphone and your room. If a tool cannot be tried before you pay, that is worth weighing too.
Reading the result
A few percentage points of WER matter less than where the errors fall. Software that misses the occasional short word but gets every name right is more pleasant to use than software with a better score that mangles your colleagues' names. Correcting a wrong word costs a moment; noticing it costs attention.
So judge the transcript, not the number. Read what came out, ask how much tidying it needs, and ask whether that is less work than typing it would have been. That is the whole question.
Where CaringDictate sits
We do not publish an accuracy percentage. We could produce one — everybody can — but it would be measured on recordings that are not your voice, and presenting it as a prediction of your experience would be misleading.
CaringDictate uses open speech-recognition models that run on your own computer, and it ships several model sizes so it can fit the machine you have. There is a free seven-day trial, which exists precisely so this question gets settled by your own voice rather than by our marketing. If the free voice typing already built into Windows scores well enough for you when you run the test above, that is a good outcome — use it and keep your money.
Where to go from here
The voice typing glossary explains the rest of the vocabulary you will meet while shopping. The Dragon comparison covers product-level differences, and the privacy guide covers where your audio goes in local versus cloud tools.