AI and conversational IVR
IVR testing with real human callers
Real people call the AI IVR or conversational IVR your customers already reach, and use it end to end the way they would. They report where it routed them, what it misheard, and whether they would have hung up before reaching the right place.
Where testers reach it
- Phone line
What you get
- Every transcript or recording
- Scores on the 8 criterion rubric
- A one-page verdict
- No fix to sell, ever
Most IVR testing checks that the call flow is wired correctly: press 2, reach billing, hear the right prompt. That kind of testing is necessary, and it is well served by automated tools. What it does not tell you is how a caller experiences the menu when they do not know which option they need.
Conversational IVR and AI IVR systems replace the keypad with open questions such as "Tell me why you are calling." That is friendlier when it works and more confusing when it does not. A caller who describes the problem in their own words can easily land in the wrong queue, get asked to repeat themselves, or loop back to the start.
Our IVR testing services use human callers with real reasons for calling. They describe problems the way customers do, change their minds, press zero, and say "operator" in several ways. Each call is scored on the fixed rubric and ends with a plain note on where the caller ended up and whether it was the right place.
We test the IVR you already have. We do not design, build or tune call flows.
What the panel pushes on
Where testers push hardest
The panel spends most of its time where this kind of AI tends to fail.
Routing by intent
Testers describe their reason for calling in their own words and record which queue or answer they reached, and whether it was the right one.
Speech recognition
Account numbers, dates of birth, names and addresses spoken naturally, with accents and at different speeds.
Menu depth and loops
Testers count the steps to reach the right place and flag every loop back to the main menu.
Escape routes
Pressing zero, saying "agent" or staying silent. Testers check whether a caller who needs a person can reach one.
Transfers and context
When the call reaches a queue or a person, testers note whether they had to repeat everything they already said.
Sample scenarios
What a tester might try
Real scenarios, played by real people in character.
Confused first-timer
Has a problem that fits two menu options and describes it in a long sentence instead of a keyword.
Asks for a human
Presses zero repeatedly, then says "operator", then "talk to someone" to find which escape route works.
Accent or non-native speaker
Reads out an account number with a regional accent and corrects one digit halfway through.
Angry customer
Calls about a service outage, says so in the first sentence and gets impatient at each extra question.
Edge case
Calls about an account that is not in their name, on behalf of an elderly relative.
You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.
The rubric
Scored on the same eight criteria as every test
Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood meFocus | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it right | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it done | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | EffortFocus | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | HandoffFocus | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in bounds | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stayFocus | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
Why people, not only software
What automated checks miss here
Automated IVR testing is excellent at checking that every path connects and every prompt plays. It follows the menu the way it was designed. Human callers do not, and the gap between the designed path and the path a real caller takes is where most IVR frustration lives.
Containment rates and average handle time can improve while callers are quietly giving up. A human tester tells you the exact prompt where they would have hung up.
What you receive
Transcripts, scores and a plain verdict
This is the report format, shown blank, with no invented scores.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Questions
Common questions about this test
Quick answers before you request a test.
Do you replace automated IVR testing tools?
No. Automated tools check that call flows are wired correctly at scale. We add the human side: whether a real caller understood the prompts, reached the right place, and would have stayed on the line.
Can you test a traditional keypad IVR as well as a conversational one?
Yes. Any IVR a customer calls can be tested. The rubric is the same, and testers note where touch tone and speech options behave differently.
How many calls does an IVR test include?
The Audit starts from $349 for a 12-session panel. Larger IVRs with many branches usually suit a Custom Panel, which we quote after a short scoping form.
Keep reading
Related guides and use cases
More on testing this kind of AI well.
After an Update
Rerun the same rubric after a change and see what moved, criterion by criterion.
Read more about After an UpdateUse caseAccessibility Testing
Testers with accessibility needs check whether your AI actually works for them.
Read more about Accessibility TestingUse caseMultilingual Testing
Native and non-native speakers test whether your AI truly works in every language.
Read more about Multilingual TestingGuideHow to Test a Voice Agent Before Go-Live
Real callers, accents, interruptions, noise and handoff.
Read more about How to Test a Voice Agent Before Go-LiveGuideChatbot KPIs and What They Miss
Containment, deflection, CSAT and the gaps between them.
Read more about Chatbot KPIs and What They MissGuideAI Agent Evaluation: Automated vs Human
Automated evals, simulations and human evaluation, explained for owners.
Read more about AI Agent Evaluation: Automated vs HumanSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.