Glossary
The language of voice computing
Short, straight definitions of the terms behind dictation, voice agents, and voice-to-action software, written to answer the question, not to sell.
What is voice to action?
Voice to action is the ability to complete tasks on a computer by speaking, rather than only converting speech into text.
What is a voice agent?
A voice agent is an AI assistant you operate by speaking, which can understand a goal, use context such as what is on your screen, and carry out multi-step tasks in software on your behalf.
What is dictation software?
Dictation software converts your live speech into written text at the cursor, in real time, so you can write by talking.
What is word error rate?
Word error rate (WER) is the standard accuracy metric for speech recognition.
What is the difference between push-to-talk and hands-free voice input?
Push-to-talk and hands-free are the two ways of telling a voice system when to listen.
What is a wake word?
A wake word is a chosen phrase, such as "Hey Siri" or "Alexa," that a voice system listens for continuously in order to know when to start processing what you say.
What is screen-aware AI?
Screen-aware AI is an assistant that can use what is currently on your display as context for a request, so you can ask "what does this error mean" or "reply to this thread" without describing or copying anything.
What is custom vocabulary in dictation?
Custom vocabulary (or a custom dictionary) is a user-maintained list of words a dictation system should recognize and spell exactly as specified: names, company terms, jargon, and invented words that generic speech recognition would otherwise mistranscribe.