Guide
How to test a chatbot with real people
A practical, step by step method for finding out how your customer-facing chatbot handles real customers, before and after it goes live.
By RealHumanTests Updated 14 minute read
If you want to know how to test a chatbot that your customers actually use, the most reliable method is also the oldest: put real people in front of it, ask them to behave like your customers, and write down exactly what happened. Automated checks confirm that the bot returns an answer. People tell you whether that answer would have kept a customer or lost one.
This guide walks through a complete chatbot test you can run yourself, whether you have a support bot on your website, a chat assistant inside your app, or a bot that answers on WhatsApp or Messenger. It is written for the person who owns the customer experience, not the engineer who built the bot. You will not need access to the model, the prompt, or the vendor dashboard. You only need the chatbot, a handful of people, and a consistent way to score what they saw.
Why test a chatbot with people
Most chatbots arrive with some form of built-in reporting: conversation counts, containment rates, thumbs up and thumbs down. Those numbers are useful, and our guide to chatbot KPIs explains what each one measures. What they cannot tell you is why a conversation ended. A customer who closes the chat window because the problem was solved and a customer who closes it because they gave up look identical in most dashboards.
Automated testing tools have the same blind spot from the other direction. A simulated user can send thousands of messages and check each reply against an expected answer. That is valuable for catching regressions at volume. But a machine grading a machine cannot tell you whether a real person, tired and a little annoyed, would have understood the reply, trusted it, and stayed. That judgment is the whole point of a customer-facing chatbot, and only a person can make it.
Human testing is not a replacement for either. It is the layer that tells you what the numbers mean.
Step 1: Write down what the chatbot is for
Before anyone types a message, write one short paragraph describing what the chatbot is supposed to do. Be specific. "Answer customer questions" is not a scope. "Answer order status, shipping and return policy questions for US customers, collect details for damaged item claims, and hand off to a person for anything involving a refund over the standard policy" is a scope.
Your scope should answer four questions:
- What tasks should a customer be able to finish without talking to a person?
- What information is it allowed to give, and where does that information come from (your help center, your product catalog, your policy pages)?
- What must it never do, such as promise refunds, quote custom prices, or give medical or legal advice?
- When should it hand off to a person, and how does that handoff work in practice?
This paragraph becomes your answer key. When a tester reports that the bot promised a free replacement, you need to know whether that was allowed. Keep a copy of your current published policies next to it so testers and scorers can check facts against the source.
Step 2: List what worries you
Every team that owns a chatbot has a private list of things they suspect it handles badly. Write that list down. Common entries include:
- Refund and cancellation requests, especially from customers who are already upset
- A new product, plan or policy that launched after the bot was set up
- Customers who ask for a person straight away
- Questions that sit just outside the bot's scope
- Pricing, discounts and promotions
- Customers who write in short fragments, with typos, or in a second language
These concerns decide where you push hardest. A good test spends most of its sessions on the areas that matter most to your business, rather than spreading evenly across every possible question.
Step 3: Choose your personas
A persona is a type of customer, described by how they behave rather than who they are. Personas keep testers from all acting like polite, patient, articulate power users, which is how most internal testing goes wrong. Start with these eight:
| Persona | What they push on |
|---|---|
| Angry customer | Tone, escalation, handoff |
| Confused first-timer | Understanding, effort |
| Off-topic wanderer | Staying in bounds, recovery |
| Refund or discount demand | Policy accuracy, unauthorized promises |
| Edge case | Unusual orders, accounts and dates |
| Asks for a human | Handoff |
| Accessibility needs | Screen readers, cognitive load, plain language |
| Accent or non-native speaker | Understanding, multilingual handling |
Add personas that are specific to your business. A subscription company might add "customer trying to cancel who does not want a retention offer". A retailer might add "gift buyer who does not have the order number". Match your personas to your worry list from step 2.
Step 4: Write scenarios, not scripts
A scenario gives the tester a goal, a situation and a mood, then lets them talk naturally. A script tells them exactly what to type. Scripts are useful for automated regression tests. For human testing they defeat the purpose, because the value of a person is that they phrase things the way real customers do.
A good scenario has four parts:
- Who you are: "You bought a jacket as a gift three weeks ago. You do not have the order number."
- What you want: "You want to exchange it for a larger size."
- How you feel: "You are in a hurry and mildly irritated that there is no phone number."
- When to stop: "Stop when the exchange is started, when you are told clearly it cannot be done, or when you would realistically give up."
That last part matters. Real customers give up. If you tell testers to keep going until they succeed, you will measure their persistence, not your chatbot. Ask them to stop where a real customer would, and to write down that moment.
Step 5: Recruit and brief testers
Your testers should not be the people who built or configured the chatbot. They know what it can do and unconsciously steer around its weak spots. Colleagues from other departments are better. People outside the company are better still, especially people who resemble your actual customers in age, comfort with technology and first language.
For a first test, aim for a panel large enough that every persona and every high priority concern gets at least one session, and your biggest concerns get two or three. A dozen well planned sessions will teach you more than a hundred unplanned ones.
Brief each tester with:
- Their scenario card, and nothing about how the chatbot works internally
- A reminder to type the way they naturally would, typos included
- An instruction to capture the full transcript (a screenshot series or the export, if available)
- The scoring sheet they will fill in immediately after the session, while it is fresh
Step 6: Run the sessions
Spread sessions across different days and times if your chatbot behaves differently out of hours (many bots offer a handoff during business hours and a contact form at night). Test on the devices your customers use: a phone browser is a very different experience from a desktop widget.
Ask testers to note, as they go, the moments where they hesitated, re-read something, or felt the urge to leave. Those notes are often more useful than the transcript itself.
Customer: need to swap a jacket i got as a gift, wrong size
Bot: I can help with returns! Please enter your order number to get started.
Customer: dont have it, it was a gift
Bot: I can help with returns! Please enter your order number to get started.
Tester note: Second identical reply. A real customer would give up here. Would I stay: 1.
Step 7: Score every session the same way
The most common reason a chatbot test produces a pile of anecdotes instead of a decision is inconsistent scoring. Use one rubric for every session, and ask testers to score immediately after they finish. We use eight criteria, each scored 1 to 5 with a written reason:
- Understood me: did it understand what I actually asked, including when I phrased it badly?
- Got it right: was every fact, price and policy correct, with nothing made up?
- Got it done: did I leave with my problem solved or my task completed?
- Effort: how many turns, repeats and rephrasings did it take?
- Handoff: when I needed a person, could I reach one without a fight?
- Tone: did it stay patient and respectful when I was angry, confused or slow?
- Stayed in bounds: did it avoid promises, discounts or advice it had no authority to give?
- Would I stay: would I have hung up, given up, or gone to a competitor?
The full rubric, with what a 1 and a 5 look like for each criterion, is published on our rubric page. Feel free to use it. The written reason is not optional: a score of 2 on Got it right means nothing until you know which fact was wrong.
After scoring, a second person should check the "Got it right" scores against your published policies. Testers can tell when an answer sounds wrong, but they cannot always know your returns window or your current price.
Step 8: Read the results
With every session scored, look at the results in this order:
- Stayed in bounds failures first. Any session where the bot promised something it should not have is a business risk, regardless of how the other scores look.
- Would I stay, by persona. If angry customers and confused first-timers score low while everyone else scores high, you know exactly which customers you are losing.
- Handoff. A chatbot that cannot solve something is fine. A chatbot that will not let the customer leave to find someone who can is not.
- Patterns in the written reasons. Group the reasons. Three testers writing "it repeated itself" is one finding, not three.
Then make a plain call. Is the chatbot good enough to keep as it is? Does it mostly work, with specific behaviors that need attention from whoever maintains it? Or did testers give up often enough that the setup deserves a serious rethink? Write that verdict down, with the transcripts that support it, and hand it to the people responsible for the bot.
When to test a chatbot
A human test is most valuable at four moments:
- Before launch, when fixing problems is cheapest and no customer has seen them. See pre-launch testing.
- After any change to the model, the vendor, the prompt, the knowledge base or your policies. See testing after an update.
- On a regular schedule, because chatbots can change behavior without anyone on your team touching them. Our guide to AI drift explains why.
- When the numbers look odd, such as a sudden jump in containment or a drop in handoffs, which can mean customers are giving up rather than being helped.
Common mistakes
- Testing with the team that built the bot, who know how to phrase things so it works.
- Writing scripts instead of scenarios, so every tester types the same polite sentence.
- Telling testers to keep going until they succeed, which hides the moment real customers give up.
- Scoring from memory days later instead of right after each session.
- Changing the rubric between tests, which makes before and after comparisons meaningless.
- Skipping the fact check, so confident wrong answers get scored as helpful.
- Testing only on desktop when most customers arrive on a phone.
- Running one test at launch and never again.
For a ready made list of scenarios to run, see the chatbot testing checklist. If you would rather have an independent panel run the test for you, RealHumanTests supplies real people who play your customers against the chatbot you already have, scores every session on this rubric, and hands back the transcripts with a plain verdict. We only diagnose; we never sell the fix. You can see how it works or read more about support chatbot testing.