Guide
How to test a voice agent before and after go-live
Voice agents fail in ways chat never does: accents, background noise, interruptions and silence. Here is how to test for all of it.
By RealHumanTests Updated 12 minute read
If you are wondering how to test a voice agent before customers start calling it, the short answer is: call it, a lot, as the people who will actually call it. An AI voice agent that sounds natural in a quiet demo can struggle with a strong accent, a barking dog, a caller who interrupts, or a caller who pauses to find their account number. None of that shows up until real people ring the line.
This guide covers voice agent testing for AI phone agents, AI receptionists and conversational IVR menus, before go-live and after. It is written for the business owner or operations lead who owns the phone line, and every step can be done without access to the model or the vendor's configuration.
Why voice agents fail differently
Chat and voice agents can share the same underlying model and knowledge, but voice adds layers that chat never has to deal with:
- Speech recognition. Before the agent can understand a caller, it has to turn their voice into words. Accents, speech differences, fast talkers, quiet talkers and noisy rooms all affect that step, and every later decision depends on it.
- Timing. People interrupt, pause, say "um", and talk over the agent. An agent that stops talking too readily sounds jumpy; one that never stops sounds rude.
- No scrollback. A caller cannot re-read what the agent said. Long lists, long numbers and long policies are hard to follow by ear.
- Patience runs out faster. Callers who feel stuck on the phone tend to hang up quickly, and some will call a competitor instead. That makes the "Would I stay" judgment especially important on voice.
Automated voice testing tools can place many synthetic calls and check responses at scale, which is useful for regressions. What synthetic voices do not give you is the range of real human speech, or a real caller's honest answer to "would I have hung up?".
Before you place a single call
Prepare three things:
- A scope. What should callers be able to do (book, reschedule, check hours, leave a message, get order status)? What must the agent never do? When should it transfer to a person, and to whom?
- An answer key. Your current hours, locations, prices, policies and booking rules, so scorers can check every fact the agent states.
- Consent and authorization. Test only a line you own or have written permission to test. If calls are recorded, make sure every party consents. Some US states require all parties to consent to call recording, so treat all-party consent as the default.
Choose callers who sound like your customers
The biggest single improvement you can make to a voice test is caller variety. If your internal team all sound alike, your test will tell you how the agent handles one kind of voice. Aim for a mix of:
- Regional accents from the areas you serve
- Non-native speakers of the language the agent uses
- Older callers and younger callers
- Fast talkers and slow, careful talkers
- Callers with speech differences, if you can recruit people willing to help
- Bilingual callers who switch languages mid sentence, if your customers do
Give each caller a persona as well as a voice: angry, confused, in a hurry, asks for a human immediately, calling on behalf of someone else. Our multilingual testing and accessibility testing pages go further on both.
Test real calling conditions
Your customers do not call from a recording studio. Ask callers to recreate the conditions your customers are actually in:
- A mobile phone on speaker in a car
- A busy kitchen, a shop floor or a street
- A weak mobile signal with dropouts
- A headset with a cheap microphone
- Children, pets or a television in the background
Also test from different phone types and carriers if you can. Note in the scorecard which conditions each call used, so you can see whether failures cluster in one environment.
Test how the conversation flows
These behaviors separate a voice agent callers tolerate from one they hang up on:
- Interrupting. Callers should cut the agent off mid sentence with a new question. Does it stop, listen and answer, or keep reading its script?
- Pausing. Callers should go quiet for several seconds, as if looking for a card or an email. Does the agent wait, prompt gently, or hang up?
- Filler and corrections. "I need an appointment on, um, Tuesday, no, Wednesday." Does the agent pick up the correction?
- Several things at once. "I want to move my appointment and also ask about parking." Does it handle both, or drop one?
- Going off script. A question it was not set up for. Does it admit it cannot help and offer a route that can, or does it invent an answer?
Bot: I can book you for Tuesday at 10 AM. Shall I confirm that?
Customer: No wait, Wednesday, sorry, Wednesday is better
Bot: Great, you are confirmed for Tuesday at 10 AM.
Tester note: Correction ignored. Caller on speaker in a car. Got it done: 1, Would I stay: 2.
Test names, numbers, dates and spelling
Voice agents take actions based on what they heard. A misheard digit can book the wrong day or look up the wrong account. Every test should include:
- Unusual names, and names that need spelling out letter by letter
- Email addresses read aloud, including dots, dashes and numbers
- Phone numbers, order numbers and postcodes, read in different rhythms
- Relative dates such as "next Friday" and "the day after tomorrow"
- Times near noon and midnight, and callers who say "half past" or "quarter to"
Check whether the agent reads important details back before acting, and whether it lets the caller correct them easily. Then verify what was actually recorded in your booking or CRM system, not just what the agent said.
Test the handoff to a person
On the phone, a caller who cannot reach a person when they need one is a caller who hangs up. Test:
- Asking for a person in the first sentence
- Asking for a person after the agent has failed twice
- Saying "operator", "representative", "agent" and "a real person"
- Asking out of hours, when nobody is available
- Whether the person who picks up already knows why the caller is calling
A good handoff is prompt, honest about wait times or availability, and passes along what the caller has already said. A bad one loops, argues, or drops the call.
Score every call the same way
Use one rubric for every call, scored right after hanging up. Our eight criteria are Understood me, Got it right, Got it done, Effort, Handoff, Tone, Stayed in bounds and Would I stay, each scored 1 to 5 with a written reason. On voice, two deserve special attention:
- Understood me captures speech recognition problems. Ask callers to note what the agent seemed to hear versus what they said.
- Would I stay captures the hang-up moment. Ask callers to note the exact point where a real customer would have ended the call, even if they kept going for the test.
The full rubric, with what a 1 and a 5 look like, is on our rubric page. Record calls if everyone has consented, so scorers can review disputed moments.
Before go-live and after
Before go-live, run a full panel covering every persona, caller type and calling condition. This is the cheapest time to learn the agent struggles with a particular accent or cannot handle rescheduling. See pre-launch testing.
After go-live, test again after any change to the voice, the speech recognition provider, the model, the prompt, your hours or your booking rules. Vendors also update models on their own schedules, so a smaller panel on a regular cadence will catch changes nobody announced. Our guide to AI drift explains why this happens.
Voice agent testing checklist
- Written scope, answer key and authorization are in place, with all-party recording consent.
- Callers cover the accents, languages, ages and speaking styles of your real customers.
- Calls include car speakerphones, noisy rooms, weak signal and background voices.
- Every caller interrupts, pauses and corrects themselves at least once.
- Names, emails, numbers and relative dates are spoken, read back and verified in your systems.
- Handoff is tested early, late, out of hours and with different wording.
- Every call is scored on the same rubric, with the hang-up moment noted.
- The panel is repeated after changes and on a schedule.
If you would like an independent panel to make these calls, RealHumanTests supplies real callers with real accents and real impatience who test the voice agent you already run, score every call on a published rubric, and return the recordings with a plain verdict. We only diagnose and never sell the fix. Read more about AI voice agent testing or AI IVR testing.