Glossary

What is word error rate?

How it is calculated

A recognized transcript is aligned against a human reference transcript. Every substituted word (said "affect", wrote "effect"), inserted word (wrote a word never spoken), and deleted word (dropped a spoken word) counts as one error. Errors divided by reference words gives WER, which can exceed 100% on garbled output.

Why quoted numbers rarely match your experience

WER depends heavily on the test set: microphone quality, background noise, accent, vocabulary, and speaking style. A benchmark on clean read speech says little about you dictating jargon on a laptop mic. Names and technical terms drive most real-world errors, which is why custom vocabulary matters more than a headline decimal.

What it misses for dictation

WER treats all errors equally, but for dictation, formatting quality matters too: punctuation, capitalization, numbers, and how filler is handled are scored poorly or not at all by raw WER. Two tools with identical WER can produce very different amounts of cleanup work, so judge dictation software by corrections you actually make.

Frequently asked questions

What is a good word error rate?

On clean single-speaker audio, modern systems commonly land in the low single digits, but the honest answer is contextual: measure on your own voice, microphone, and vocabulary. The number that matters is how many corrections per paragraph you personally make.

Does lower WER always mean better dictation?

Not by itself. Punctuation, formatting, filler handling, latency, and vocabulary adaptation are outside raw WER but dominate how usable dictated text is.

Keep exploring

Hear it in practice

VoiceOS puts dictation, editing, and voice-to-action in one layer on Mac and Windows.

Try VoiceOS