Support chatbots
Chatbot testing by real people who play your customers
You already have a support chatbot, or you are about to launch one. Our testers chat with it as your customers would, push on the areas that worry you, and report exactly what it did, scored on a fixed human rubric.
Where testers reach it
- Website chat widget
- In-app chat
- Facebook Messenger
- Instagram Direct
What you get
- Every transcript or recording
- Scores on the 8 criterion rubric
- A one-page verdict
- No fix to sell, ever
Chatbot testing usually starts with a list of questions someone on the team typed in before launch. The bot answered them, the answers looked fine, and it went live. The trouble is that your customers never type those questions. They paste an order number into the wrong field, describe the problem in three messages instead of one, get annoyed halfway through, and ask for a person the moment they feel stuck.
A real chatbot test puts people in front of the bot who behave like that. Each tester plays a customer with a specific situation and a specific mood, follows the conversation wherever the bot takes it, and writes down what happened in plain words. You end up with transcripts that show how the bot handles the conversations it will actually get, not the ones it was built for.
You can tell us what worries you before testing starts. A new returns policy, a product line the bot was never trained on, a spike in complaints about the chat widget: we write scenarios around it. Every session is still scored on the same fixed rubric, so the results are comparable across runs and across months.
This is AI chatbot testing, not support outsourcing. We never answer your customers and we never touch your bot. We use it the way your customers do, then tell you what we saw.
What the panel pushes on
Where testers push hardest
The panel spends most of its time where this kind of AI tends to fail.
Policy and fact accuracy
Testers ask about returns windows, fees, shipping times and eligibility, then check every answer against your own published policy. A confident wrong answer is logged word for word.
Messy, real phrasing
Typos, half sentences, two questions in one message, and the question asked a second way after the first answer missed. This is where many bots quietly fail.
Reaching a person
Testers ask for a human early, late, politely and angrily, and record whether the bot hands off, how long it takes, and whether the context comes along.
Frustration and tone
An angry customer tests whether the bot stays calm and plain, or turns robotic, repetitive or preachy at the worst moment.
Staying in bounds
Testers push for refunds, discounts and exceptions to see whether the bot promises something your team would never approve.
Dead ends and loops
Every time a conversation circles back to the same menu or answer, the tester notes it and says whether they would have given up at that point.
Sample scenarios
What a tester might try
Real scenarios, played by real people in character.
Refund or discount demand
Asks for a refund on an item bought 45 days ago when the policy says 30, then asks for store credit instead, then asks for a discount on the next order.
Confused first-timer
Cannot find their order number, describes the product by color instead of name, and is not sure whether they have an account.
Angry customer
Third contact about the same late delivery, opens with a complaint in capitals and asks why nobody has fixed it yet.
Asks for a human
Types "agent" as the first message, then "real person please" when the bot offers a menu instead.
Off-topic wanderer
Starts with a billing question, drifts into asking for product recommendations and then for opinions on a competitor.
Accent or non-native speaker
Writes in simple English with grammar from another language and switches to Spanish for one message when the bot misunderstands.
You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.
The rubric
Scored on the same eight criteria as every test
Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood meFocus | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it rightFocus | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it done | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | Effort | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | HandoffFocus | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in bounds | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stayFocus | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
Why people, not only software
What automated checks miss here
Automated chatbot testing tools are good at volume. They can send thousands of scripted or simulated messages and flag answers that break a rule. What they cannot tell you is whether a real person, reading that answer on a phone while annoyed, would have trusted it, understood it, or closed the tab.
Vendor dashboards report on the conversations the bot thinks went well. A customer who gave up and called instead often counts as a contained conversation. A human tester reports the moment they would have left, and why, in their own words.
What you receive
Transcripts, scores and a plain verdict
This is the report format, shown blank, with no invented scores.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Questions
Common questions about this test
Quick answers before you request a test.
Do you need access to our chatbot's settings or training data?
No. Testers use the chatbot exactly where your customers do, on your website, app or messaging channel. We only need to know where it lives, what it is meant to handle, and any test accounts or orders you want used.
Can you test a chatbot that is not live yet?
Yes. If it runs on a staging site or a private link, testers can use that. Pre-launch testing is one of the most common reasons companies bring us in.
How is this different from the test questions our team already ran?
Your team knows what the bot is supposed to say and tends to ask the way it expects. Our testers do not. They arrive with a customer's situation and mood, follow the conversation wherever it goes, and score it on a rubric that includes whether they would have stayed.
Keep reading
Related guides and use cases
More on testing this kind of AI well.
Pre-Launch Testing
A human panel runs your AI through real customer behavior before it goes live.
Read more about Pre-Launch TestingUse caseAfter an Update
Rerun the same rubric after a change and see what moved, criterion by criterion.
Read more about After an UpdateUse caseHuman Red Teaming
Real people try to talk your AI into promises, prices and answers it should refuse.
Read more about Human Red TeamingGuideHow to Test a Chatbot With Real People
Scope, personas, scenarios, scoring and reading the results, step by step.
Read more about How to Test a Chatbot With Real PeopleGuideThe Chatbot Testing Checklist
A printable list of scenarios and personas for any chatbot.
Read more about The Chatbot Testing ChecklistGuideChatbot Failures: Real Cases and Lessons
Documented public failures and the human test that catches each.
Read more about Chatbot Failures: Real Cases and LessonsBenchmarkSupport Chatbot Index
Website and in-app chatbots that answer customer support questions.
Read more about Support Chatbot IndexSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.