The speed-vs-empathy trap in ticket replies
Every support agent hits the same wall: typing a genuinely thoughtful reply takes long enough that it's tempting to either keep it short or lean on a saved macro. Neither serves the customer well. The gap is bigger than it feels: people speak in natural conversation at roughly 120-150 words per minute, while typing on a physical keyboard tops out around 35-65 words per minute for most people, with a large 2019 study across three universities putting two-thumb phone typing at about 38 words per minute. That 3-4x gap is exactly why a rushed reply often gets a template pasted into it, and a considered one takes minutes an agent doesn't have during a busy queue.[1] [2]
- Natural conversational speech: about 120-150 words per minute (National Center for Voice and Speech data, cited via VirtualSpeech).
- Typical keyboard typing: about 35-65 words per minute; two-thumb phone typing averages 38 wpm (Aalto University, University of Cambridge and ETH Zurich, MobileHCI 2019, 37,000+ participants).
- 82% of customers say they'd still rather deal with a human than a bot, even if the wait time and outcome are identical (HubSpot/SurveyMonkey survey, reported by Customer Experience Dive, Aug 2025).
- 86% say empathy and human connection matter more to them than a fast reply (Five9 survey, reported by Customer Experience Dive, Aug 2025).
Why the fastest reply isn't always the safest one
The instinct to answer a busy queue with saved macros makes sense: typing takes time, and a template is fast. But a copy-pasted answer is easiest to spot in exactly the moment it hurts most, when a customer is already frustrated and wants to feel heard rather than processed. Zendesk's 2026 CX Trends report, based on more than 11,000 consumers and business leaders across 22 countries, found that 75% of consumers are fine with agents using AI to help draft a response. That's a real, cited number about willingness, not a claim about whether anyone, human or automated, could detect the difference afterward; this guide doesn't measure or claim anything about AI-detection tools. What consumers are telling researchers is closer to a preference: most still say they'd rather deal with a human than a bot even when the wait time and outcome are identical, and most say empathy matters more to them than getting an answer fast. Consumers aren't rejecting speed itself. They're rejecting the feeling of being handled instead of helped, which is a different problem than word count, or detectability, can fix.[3] [4]
Voice typing your reply is not the same as letting AI write it
This is the distinction that matters most for this workflow, and it's easy to blur, so it's worth stating plainly: an AI-drafted reply tool reads the ticket and generates the wording for the agent to review, edit, or send, the phrasing, structure, and empathetic language come from the model, not the person named on the reply. VoiceOS's Dictation Mode does the opposite. The agent composes the reply out loud, in their own words, deciding what to say and how to say it before a single sentence exists in text form, and the software's job is limited to turning that speech into clean text. Depending on the Polishing setting on the Dictation tab, it removes filler words and false starts, fixes punctuation and formatting, and, only at the highest 'Polished' tier, smooths awkward phrasing. It does not invent sentences, decide what to say, or generate empathy on the agent's behalf. This isn't a claim about whether a customer, or an AI-detection tool, could tell a dictated reply apart from a typed one after the fact; that's not something this guide measures, and it isn't the point. The point is upstream of detectability: who decided what the customer reads. What the customer reads is still the agent's own judgment call about how to handle their specific issue, just captured at speaking speed instead of typing speed.[5]
| AI-drafted reply | Voice-typed reply (VoiceOS Dictation) | |
|---|---|---|
| Who chooses the wording | The AI model, from the ticket content | The agent, spoken in their own words |
| What the software does | Generates new sentences and structure | Transcribes speech; optionally cleans up filler, punctuation, and (on Polished) awkward phrasing |
| Agent's role | Reviews and edits AI output before sending | Composes the reply in real time by speaking it |
| Where the speed comes from | Generation is near-instant, but reviewing and editing still takes time | Speaking is roughly 3-4x faster than typing, so composing itself is faster |
The five-step workflow: from open ticket to sent reply
Dictation Mode works anywhere there's a text cursor, including Slack, Gmail, Notion, and browser fields generally, which covers the reply box in Zendesk, Intercom, Freshdesk, or whatever helpdesk a team already uses, without a separate integration to set up. The mechanics stay the same for every ticket; only what you say changes. Here's the workflow, step by step:[5]
- Step 1 — Read before you speak: skim the ticket once for what the customer actually needs, plus any specifics you'll reference by name, an order number, an error message, a dollar amount, a plan tier.
- Step 2 — Open the reply and start dictating: click into the reply field and hold the dictation shortcut (Fn on macOS, Ctrl+Shift on Windows) to start talking.
- Step 3 — Speak the whole reply out loud: greeting, the specific answer or fix, and the next step, using the customer's name and the specifics from step 1 so it reads as theirs and not a template.
- Step 4 — Release and read it back: check that dictated numbers landed correctly before sending; misheard digits in an order number, a refund amount, or a date are the most common dictation error, and they matter more in a support reply than almost anywhere else.
- Step 5 — Send, and bank the boilerplate: if part of the reply was the same line you say on most tickets, a refund-policy note, a sign-off, save it as a Replacement so the next ticket needs less speaking, not more typing.
The specifics change by ticket type, but the shape doesn't. Two illustrative examples, a billing issue and a technical issue, written for this guide to show the shape of a spoken reply, not transcribed from a real ticket:
Illustrative dictated replies by ticket type (not real customer tickets)
| Ticket type | What the agent sees | What the agent might dictate |
|---|---|---|
| Billing | Customer reports being charged twice for their invoice. | "Hi Maria, thanks for flagging that. I can see your invoice was charged twice, so I've refunded the duplicate charge to the card on file. It usually takes a few business days to show back up on your statement. Let me know if it hasn't landed and I'll take another look." |
| Technical | Customer reports a sync error after a recent update. | "Hi Sam, thanks for sending that error message. That usually happens when one of your devices is still on an older version. Can you check that everything is updated and try syncing again? If it still fails after that, reply here with a screenshot and I'll look into your account directly." |
Staying fast without turning into a macro
The features that make Dictation Mode useful for support work aren't the transcription itself, they're the personalization tools sitting next to it on the Dictation tab, including the boilerplate step from the workflow above. Replacements let an agent save short spoken phrases that expand into a saved snippet, like a return-policy link or an email sign-off, up to 100 of them; the agent still says the empathetic opening and the specifics out loud, and only the boilerplate part is saved. Custom Prompts add per-app writing guidance, so the tone used in a Zendesk reply can differ from the tone used in an internal Slack note, without the agent having to remember to switch styles by hand. Spelling makes sure a customer's name or a product's actual name comes out right instead of being autocorrected into something generic, a small thing that matters more in a support reply than almost anywhere else.[5]
Privacy and accuracy limitations
Ticket replies often include a customer's name, order number, account email, or billing details, so it's worth knowing what happens to that audio. By default, VoiceOS does not use your raw audio, transcripts, or edited text to train its models, and raw audio and transcripts are deleted immediately after the processing needed to produce your text, unless the optional 'Help Improve VoiceOS' setting is turned on, in which case that content may be retained indefinitely for evaluation. Producing the transcript at all does require sending your audio to a speech-processing provider; that transmission is encrypted in transit (TLS 1.2+) and any stored data is encrypted at rest (AES-256). None of this replaces your own company's policy for customer information: treat a dictated reply with the same care as a typed one, and check your organization's rules on where customer details can be spoken aloud, especially in shared workspaces.[6] [7]
On accuracy: speech-recognition performance varies with accent, background noise, and microphone quality, and this guide hasn't benchmarked any of that for a support-specific vocabulary of order numbers, product names, or error codes. Your Polishing setting also changes how closely the transcript matches what you said: 'None' preserves your phrasing word for word, while 'Polished' smooths phrasing enough that it's worth the same read-back Step 4 above calls out, especially on anything with a dollar figure, date, or account number. Treat the overall speed argument as directional, not a guarantee for your queue; actual gains depend on how much back-and-forth editing a reply needs regardless of how it was produced.[1] [2] [5]
Try it on your next shift
The easiest test is a small one: pick five tickets that would normally get a saved macro, and run the five-step workflow above on each instead, dictating a real, specific reply at whatever length the situation actually calls for. Time it against your usual typing pace, and read the reply back to check it still sounds like something you'd say out loud to that customer. VoiceOS starts with a 7-day free trial that requires a card but charges nothing until the trial ends, so there's no cost to running that comparison on your own queue.[5]
Sources and methodology
Read the VoiceOS product knowledge base, the Dictation and Agent feature pages, pricing page, Privacy Policy, and Security page to confirm product, data-handling, and encryption facts and avoid cannibalizing existing articles. Searched for published data on typing speed (Aalto University / University of Cambridge / ETH Zurich, MobileHCI 2019), conversational speaking rate (National Center for Voice and Speech, via VirtualSpeech), and consumer sentiment toward AI in customer service (Zendesk CX Trends 2026; HubSpot/SurveyMonkey and Five9 surveys reported by Customer Experience Dive). No first-party dictation testing was performed; the billing and technical reply examples are illustrative sample phrasing written for this guide, not transcribed from real tickets. Limitations: No first-party test compared dictated, typed, and AI-drafted replies under identical conditions. Typing- and speaking-speed figures are general population averages, not agent-specific measurements. Consumer AI-sentiment stats describe attitudes toward AI in service broadly, not a controlled dictation-vs-AI-drafting test. Recognition accuracy varies by accent, noise, and hardware, unbenchmarked here. The billing and technical examples are illustrative sample phrasing, not real tickets or a first-party transcript. The privacy section summarizes VoiceOS's public Privacy Policy and Security page as of 2026-08-14, not legal or compliance advice.
- Smartphone typing speeds catching up with keyboards — Aalto University
- Average Speaking Rate and Words per Minute — VirtualSpeech
- 59 AI customer service statistics for 2026 — Zendesk
- Customers still don't love AI in customer service — Customer Experience Dive
- VoiceOS Dictation feature page — VoiceOS
- VoiceOS Privacy Policy — VoiceOS
- VoiceOS Security — VoiceOS
Frequently asked questions
Does VoiceOS write my customer support replies for me?
No. Dictation Mode only turns your spoken words into text; depending on your Polishing setting it can clean up filler words, punctuation, and (on the 'Polished' tier) awkward phrasing, but it does not generate the reply's content or wording. That's the core difference from an AI-drafting tool, which composes the reply itself from the ticket.[5]
Will voice typing work inside Zendesk, Intercom, Freshdesk, or Gmail?
Yes. Dictation Mode works in any app or browser field with a text cursor, so it works in whatever helpdesk or inbox a team already uses without a separate integration to set up.[5]
How much faster is speaking a reply than typing one?
Conversational speech runs about 120-150 words per minute versus roughly 35-65 words per minute for typical keyboard typing, a gap of roughly 3-4x. Those are general population averages, not a measurement of support-ticket replies specifically, so treat it as a directional advantage rather than a guaranteed number for any one queue.[2] [1]
Is a dictated reply "AI-written," and can AI detectors tell the difference?
No, and this guide doesn't make a claim either way about AI-detection tools; that's not something we've tested or measured. What actually differs is upstream of detectability: with Dictation Mode, the agent decides every word by speaking it themselves, and the software only transcribes (and, depending on Polishing, lightly cleans up) that speech. An AI-drafting tool decides the wording itself from the ticket, and the agent reviews it after the fact. Those are two different authorship processes, whether or not any tool could tell them apart afterward.[5]
Is it safe to dictate replies that include a customer's billing or account details?
By default, VoiceOS does not use your raw audio, transcripts, or edited text to train its models, and that content is deleted immediately after the processing needed to produce your text, unless you've turned on the optional 'Help Improve VoiceOS' setting, which retains it for evaluation. Producing the transcript does require sending your audio to a speech-processing provider, encrypted in transit and at rest. That's VoiceOS's documented handling, not a substitute for your own company's policy on where customer information can be spoken aloud.[6] [7]
What if I don't want VoiceOS changing my wording at all?
Set Polishing to 'None' on the Dictation tab for a raw, word-for-word transcript. The default 'Light' setting removes filler words and false starts and fixes punctuation without changing your phrasing; only the 'Polished' tier smooths awkward phrasing.[5]
Start Dictating Faster Ticket Replies
Try VoiceOS free for 7 days (card required, nothing charged until the trial ends) and dictate your next batch of ticket replies in your own words, at speaking speed.
Start free trial