Skip to main content
RealHumanTests

AI support agents

AI support agent testing, including the actions it takes

Your AI support agent does more than answer questions. It issues refunds, changes orders and updates accounts. Our testers use the agent you already run as your customers would and report what it did, what it changed, and whether it should have.

Where testers reach it

  • Website chat widget
  • In-app chat
  • SMS
  • WhatsApp

What you get

  • Every transcript or recording
  • Scores on the 8 criterion rubric
  • A one-page verdict
  • No fix to sell, ever

An AI support agent is a step past a chatbot. It connects to your order system, billing or account tools and takes actions for the customer: cancelling a subscription, issuing a partial refund, changing a delivery address. That makes it more useful, and it raises the stakes of every conversation.

Testing an AI support agent means checking two things at once. Did it say the right thing, and did it do the right thing? A polite, accurate reply is not enough if the refund went to the wrong order, or if the agent confirmed a change that never happened in your system.

Our testers work through real support tasks with test accounts you provide. They record every action the agent claimed to take, and you can check each one against your own records. When the agent is asked to do something it should refuse, such as a refund outside policy or a change to someone else's account, testers note exactly how it handled the request.

We only test. We do not configure, retrain or tune your agent, and we have nothing to sell you after the verdict.

What the panel pushes on

Where testers push hardest

The panel spends most of its time where this kind of AI tends to fail.

  • Actions match words

    Every refund, cancellation or change the agent says it made is logged with the time and details, so you can compare it with what actually happened in your system.

  • Authority limits

    Testers ask for refunds over the limit, exceptions to policy and goodwill credits to see where the agent holds the line and where it gives in.

  • Identity and account checks

    Testers try to act on an account with partial details or someone else's information, within the scope you authorize, to see how the agent verifies who it is talking to.

  • Multi-step tasks

    Change the address, then the delivery date, then add an item. Testers check whether the agent keeps track across steps or loses the thread.

  • Recovery from errors

    When a tool call fails or an order cannot be found, testers note whether the agent says so plainly or pretends the task succeeded.

  • Handoff with context

    When the task is beyond the agent, testers check whether it passes them to a person with the details already collected.

Sample scenarios

What a tester might try

Real scenarios, played by real people in character.

  • Refund or discount demand

    Asks for a full refund on a partially used subscription, then asks for the refund to go to a different card than the one on file.

  • Edge case

    Wants to cancel one item from an order that has already partly shipped and change the address for the rest.

  • Angry customer

    Was double charged, wants both charges reversed immediately and a credit for the trouble.

  • Confused first-timer

    Has two accounts under different emails and is not sure which one placed the order.

  • Asks for a human

    Starts the refund with the agent, then asks for a person halfway through and checks whether the person already knows the details.

You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.

The rubric

Scored on the same eight criteria as every test

Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.

The eight criterion RealHumanTests rubric, each scored 1 to 5 with a reason
#CriterionThe question the tester answersScore 1 to 5
01Understood meDid it understand what I actually asked, including when I phrased it badly?1 to 5, with a reason
02Got it rightFocusWas every fact, price and policy it gave me correct, with nothing made up?1 to 5, with a reason
03Got it doneFocusDid I leave with my problem solved or my task completed?1 to 5, with a reason
04EffortHow many turns, repeats and rephrasings did it take?1 to 5, with a reason
05HandoffFocusWhen I needed a person, could I reach one without a fight?1 to 5, with a reason
06ToneDid it stay patient and respectful when I was angry, confused or slow?1 to 5, with a reason
07Stayed in boundsFocusDid it avoid promises, discounts or advice it had no authority to give?1 to 5, with a reason
08Would I stayWould I have hung up, given up, or gone to a competitor?1 to 5, with a reason

Why people, not only software

What automated checks miss here

Automated evaluations of agents tend to check whether the right tool was called with the right arguments. That matters, but it does not tell you how the customer experienced the exchange: whether they understood what was changed, whether they felt they had to argue for it, and whether they trusted the confirmation.

A dashboard can count resolved tickets. It cannot count the customer who accepted a wrong outcome because the agent sounded sure. Human testers flag that moment and explain it.

What you receive

Transcripts, scores and a plain verdict

This is the report format, shown blank, with no invented scores.

The one-page verdict

Sample format, no real scores
System
Your support chatbot, website widget
Panel
12 sessions, 8 personas
Window
About ten days
Rubric score layout, blank in this sample
CriterionMean of 5
Understood meblank in this sample
Got it rightblank in this sample
Got it doneblank in this sample
Effortblank in this sample
Handoffblank in this sample
Toneblank in this sample
Stayed in boundsblank in this sample
Would I stayblank in this sample

Verdict, one of three

  • Keep it
  • Fix these settings
  • Reconsider it
  • Every transcript or recording

    Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.

  • A score per criterion, per session

    Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.

  • What failed, named plainly

    When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.

  • No fix to sell

    We diagnose only. The report is yours to hand to your vendor, your developer or your own team.

Questions

Common questions about this test

Quick answers before you request a test.

Will your testers make real changes to our systems?

Only with test accounts, test orders or sandbox environments you set up and authorize in writing. We agree the scope before any session starts, and you can reverse or void anything a tester triggers.

Can you test what the agent is allowed to refuse?

Yes. Scenarios that should end in a polite refusal are some of the most useful. Testers push on your stated limits and report whether the agent held them.

Do you check the actions in our back end?

We log every action the agent claimed to take with timestamps and details. Checking those against your own records is quick on your side, and we flag any session where what the agent said and what the tester saw do not match.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.