Skip to main content
RealHumanTests

Support chatbots

Chatbot testing by real people who play your customers

You already have a support chatbot, or you are about to launch one. Our testers chat with it as your customers would, push on the areas that worry you, and report exactly what it did, scored on a fixed human rubric.

Where testers reach it

  • Website chat widget
  • In-app chat
  • Facebook Messenger
  • Instagram Direct

What you get

  • Every transcript or recording
  • Scores on the 8 criterion rubric
  • A one-page verdict
  • No fix to sell, ever

Chatbot testing usually starts with a list of questions someone on the team typed in before launch. The bot answered them, the answers looked fine, and it went live. The trouble is that your customers never type those questions. They paste an order number into the wrong field, describe the problem in three messages instead of one, get annoyed halfway through, and ask for a person the moment they feel stuck.

A real chatbot test puts people in front of the bot who behave like that. Each tester plays a customer with a specific situation and a specific mood, follows the conversation wherever the bot takes it, and writes down what happened in plain words. You end up with transcripts that show how the bot handles the conversations it will actually get, not the ones it was built for.

You can tell us what worries you before testing starts. A new returns policy, a product line the bot was never trained on, a spike in complaints about the chat widget: we write scenarios around it. Every session is still scored on the same fixed rubric, so the results are comparable across runs and across months.

This is AI chatbot testing, not support outsourcing. We never answer your customers and we never touch your bot. We use it the way your customers do, then tell you what we saw.

What the panel pushes on

Where testers push hardest

The panel spends most of its time where this kind of AI tends to fail.

  • Policy and fact accuracy

    Testers ask about returns windows, fees, shipping times and eligibility, then check every answer against your own published policy. A confident wrong answer is logged word for word.

  • Messy, real phrasing

    Typos, half sentences, two questions in one message, and the question asked a second way after the first answer missed. This is where many bots quietly fail.

  • Reaching a person

    Testers ask for a human early, late, politely and angrily, and record whether the bot hands off, how long it takes, and whether the context comes along.

  • Frustration and tone

    An angry customer tests whether the bot stays calm and plain, or turns robotic, repetitive or preachy at the worst moment.

  • Staying in bounds

    Testers push for refunds, discounts and exceptions to see whether the bot promises something your team would never approve.

  • Dead ends and loops

    Every time a conversation circles back to the same menu or answer, the tester notes it and says whether they would have given up at that point.

Sample scenarios

What a tester might try

Real scenarios, played by real people in character.

  • Refund or discount demand

    Asks for a refund on an item bought 45 days ago when the policy says 30, then asks for store credit instead, then asks for a discount on the next order.

  • Confused first-timer

    Cannot find their order number, describes the product by color instead of name, and is not sure whether they have an account.

  • Angry customer

    Third contact about the same late delivery, opens with a complaint in capitals and asks why nobody has fixed it yet.

  • Asks for a human

    Types "agent" as the first message, then "real person please" when the bot offers a menu instead.

  • Off-topic wanderer

    Starts with a billing question, drifts into asking for product recommendations and then for opinions on a competitor.

  • Accent or non-native speaker

    Writes in simple English with grammar from another language and switches to Spanish for one message when the bot misunderstands.

You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.

The rubric

Scored on the same eight criteria as every test

Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.

The eight criterion RealHumanTests rubric, each scored 1 to 5 with a reason
#CriterionThe question the tester answersScore 1 to 5
01Understood meFocusDid it understand what I actually asked, including when I phrased it badly?1 to 5, with a reason
02Got it rightFocusWas every fact, price and policy it gave me correct, with nothing made up?1 to 5, with a reason
03Got it doneDid I leave with my problem solved or my task completed?1 to 5, with a reason
04EffortHow many turns, repeats and rephrasings did it take?1 to 5, with a reason
05HandoffFocusWhen I needed a person, could I reach one without a fight?1 to 5, with a reason
06ToneDid it stay patient and respectful when I was angry, confused or slow?1 to 5, with a reason
07Stayed in boundsDid it avoid promises, discounts or advice it had no authority to give?1 to 5, with a reason
08Would I stayFocusWould I have hung up, given up, or gone to a competitor?1 to 5, with a reason

Why people, not only software

What automated checks miss here

Automated chatbot testing tools are good at volume. They can send thousands of scripted or simulated messages and flag answers that break a rule. What they cannot tell you is whether a real person, reading that answer on a phone while annoyed, would have trusted it, understood it, or closed the tab.

Vendor dashboards report on the conversations the bot thinks went well. A customer who gave up and called instead often counts as a contained conversation. A human tester reports the moment they would have left, and why, in their own words.

What you receive

Transcripts, scores and a plain verdict

This is the report format, shown blank, with no invented scores.

The one-page verdict

Sample format, no real scores
System
Your support chatbot, website widget
Panel
12 sessions, 8 personas
Window
About ten days
Rubric score layout, blank in this sample
CriterionMean of 5
Understood meblank in this sample
Got it rightblank in this sample
Got it doneblank in this sample
Effortblank in this sample
Handoffblank in this sample
Toneblank in this sample
Stayed in boundsblank in this sample
Would I stayblank in this sample

Verdict, one of three

  • Keep it
  • Fix these settings
  • Reconsider it
  • Every transcript or recording

    Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.

  • A score per criterion, per session

    Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.

  • What failed, named plainly

    When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.

  • No fix to sell

    We diagnose only. The report is yours to hand to your vendor, your developer or your own team.

Questions

Common questions about this test

Quick answers before you request a test.

Do you need access to our chatbot's settings or training data?

No. Testers use the chatbot exactly where your customers do, on your website, app or messaging channel. We only need to know where it lives, what it is meant to handle, and any test accounts or orders you want used.

Can you test a chatbot that is not live yet?

Yes. If it runs on a staging site or a private link, testers can use that. Pre-launch testing is one of the most common reasons companies bring us in.

How is this different from the test questions our team already ran?

Your team knows what the bot is supposed to say and tends to ask the way it expects. Our testers do not. They arrive with a customer's situation and mood, follow the conversation wherever it goes, and score it on a rubric that includes whether they would have stayed.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.