AI support agents
AI support agent testing, including the actions it takes
Your AI support agent does more than answer questions. It issues refunds, changes orders and updates accounts. Our testers use the agent you already run as your customers would and report what it did, what it changed, and whether it should have.
Where testers reach it
- Website chat widget
- In-app chat
- SMS
What you get
- Every transcript or recording
- Scores on the 8 criterion rubric
- A one-page verdict
- No fix to sell, ever
An AI support agent is a step past a chatbot. It connects to your order system, billing or account tools and takes actions for the customer: cancelling a subscription, issuing a partial refund, changing a delivery address. That makes it more useful, and it raises the stakes of every conversation.
Testing an AI support agent means checking two things at once. Did it say the right thing, and did it do the right thing? A polite, accurate reply is not enough if the refund went to the wrong order, or if the agent confirmed a change that never happened in your system.
Our testers work through real support tasks with test accounts you provide. They record every action the agent claimed to take, and you can check each one against your own records. When the agent is asked to do something it should refuse, such as a refund outside policy or a change to someone else's account, testers note exactly how it handled the request.
We only test. We do not configure, retrain or tune your agent, and we have nothing to sell you after the verdict.
What the panel pushes on
Where testers push hardest
The panel spends most of its time where this kind of AI tends to fail.
Actions match words
Every refund, cancellation or change the agent says it made is logged with the time and details, so you can compare it with what actually happened in your system.
Authority limits
Testers ask for refunds over the limit, exceptions to policy and goodwill credits to see where the agent holds the line and where it gives in.
Identity and account checks
Testers try to act on an account with partial details or someone else's information, within the scope you authorize, to see how the agent verifies who it is talking to.
Multi-step tasks
Change the address, then the delivery date, then add an item. Testers check whether the agent keeps track across steps or loses the thread.
Recovery from errors
When a tool call fails or an order cannot be found, testers note whether the agent says so plainly or pretends the task succeeded.
Handoff with context
When the task is beyond the agent, testers check whether it passes them to a person with the details already collected.
Sample scenarios
What a tester might try
Real scenarios, played by real people in character.
Refund or discount demand
Asks for a full refund on a partially used subscription, then asks for the refund to go to a different card than the one on file.
Edge case
Wants to cancel one item from an order that has already partly shipped and change the address for the rest.
Angry customer
Was double charged, wants both charges reversed immediately and a credit for the trouble.
Confused first-timer
Has two accounts under different emails and is not sure which one placed the order.
Asks for a human
Starts the refund with the agent, then asks for a person halfway through and checks whether the person already knows the details.
You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.
The rubric
Scored on the same eight criteria as every test
Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood me | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it rightFocus | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it doneFocus | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | Effort | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | HandoffFocus | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in boundsFocus | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stay | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
Why people, not only software
What automated checks miss here
Automated evaluations of agents tend to check whether the right tool was called with the right arguments. That matters, but it does not tell you how the customer experienced the exchange: whether they understood what was changed, whether they felt they had to argue for it, and whether they trusted the confirmation.
A dashboard can count resolved tickets. It cannot count the customer who accepted a wrong outcome because the agent sounded sure. Human testers flag that moment and explain it.
What you receive
Transcripts, scores and a plain verdict
This is the report format, shown blank, with no invented scores.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Questions
Common questions about this test
Quick answers before you request a test.
Will your testers make real changes to our systems?
Only with test accounts, test orders or sandbox environments you set up and authorize in writing. We agree the scope before any session starts, and you can reverse or void anything a tester triggers.
Can you test what the agent is allowed to refuse?
Yes. Scenarios that should end in a polite refusal are some of the most useful. Testers push on your stated limits and report whether the agent held them.
Do you check the actions in our back end?
We log every action the agent claimed to take with timestamps and details. Checking those against your own records is quick on your side, and we flag any session where what the agent said and what the tester saw do not match.
Keep reading
Related guides and use cases
More on testing this kind of AI well.
Pre-Launch Testing
A human panel runs your AI through real customer behavior before it goes live.
Read more about Pre-Launch TestingUse caseHuman Red Teaming
Real people try to talk your AI into promises, prices and answers it should refuse.
Read more about Human Red TeamingUse caseAfter an Update
Rerun the same rubric after a change and see what moved, criterion by criterion.
Read more about After an UpdateGuideHow to Test AI Agents Customers Talk To
Testing agents that take actions, across chat, voice and text.
Read more about How to Test AI Agents Customers Talk ToGuideAI Agent Evaluation: Automated vs Human
Automated evals, simulations and human evaluation, explained for owners.
Read more about AI Agent Evaluation: Automated vs HumanGuideLLM as a Judge vs Human Evaluation
What a model grading a model can and cannot tell you.
Read more about LLM as a Judge vs Human EvaluationSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.