SMS and WhatsApp bots
SMS chatbot and WhatsApp bot testing by real people
Once your SMS chatbot or WhatsApp bot is set up, the next step is finding out how it handles real people. Our testers text the bot you already run the way your customers do and report exactly what it said, what it did, and where they would have stopped replying.
Where testers reach it
- SMS
What you get
- Every transcript or recording
- Scores on the 8 criterion rubric
- A one-page verdict
- No fix to sell, ever
Choosing an SMS chatbot or WhatsApp chatbot platform is the easy part. The harder question comes after setup: does it work for the people who will text it? Text conversations look nothing like website chat. Replies are short, spread out over hours, full of abbreviations, and often sent in answer to a message the bot sent days earlier.
That is why testing is the step after buying. Our testers use their own phones to text your number or WhatsApp account. They reply late, reply "yes" to the wrong question, send a photo, text STOP and START again, and ask something unrelated in the middle of a booking.
Each thread is captured and scored on the same fixed rubric we use for every AI type, with a written reason from the tester. You see how the bot behaves in the channel where your customers actually are.
We only test your AI SMS or WhatsApp bot. We never send messages to your customers and never change your setup.
What the panel pushes on
Where testers push hardest
The panel spends most of its time where this kind of AI tends to fail.
Short and ambiguous replies
"Yes", "ok", "2", a thumbs up. Testers check whether the bot understands which question the reply answers.
Long gaps
Testers come back hours later, or the next day, and see whether the bot remembers the conversation or starts over.
Opt out and consent
STOP, UNSUBSCRIBE and variations, followed by a later message, to see whether opt out is honored and how the bot responds.
Links and media
Testers follow links the bot sends, send photos and voice notes, and report what happens when the bot cannot read them.
Tone on a small screen
Long walls of text, too many messages at once, or replies that feel pushy all get flagged by people reading on a phone.
Handoff by text
Testers ask for a person and record whether one replies, how long it takes, and whether they are told what to expect.
Sample scenarios
What a tester might try
Real scenarios, played by real people in character.
Confused first-timer
Replies "what is this" to the first message, then answers a yes or no question with a question.
Edge case
Starts rescheduling, stops replying for six hours, then comes back with a different date than the one discussed.
Off-topic wanderer
Replies to an appointment reminder with a billing question and a photo of a receipt.
Asks for a human
Texts "can a real person call me" and waits to see whether anyone does.
Accent or non-native speaker
Texts in Spanish after the first English message, then mixes both languages.
You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.
The rubric
Scored on the same eight criteria as every test
Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood meFocus | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it right | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it doneFocus | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | EffortFocus | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | Handoff | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | ToneFocus | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in bounds | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stay | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
Why people, not only software
What automated checks miss here
Automated tests send clean, complete messages on a fixed schedule. Real customers send "k" at midnight in reply to something from Tuesday. People catch the moments where the bot loses the thread because they are the ones sending the messy replies.
Delivery and response rates tell you messages went out and came back. They do not tell you whether the person on the other end felt helped or pestered. A tester's written reason does.
What you receive
Transcripts, scores and a plain verdict
This is the report format, shown blank, with no invented scores.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Questions
Common questions about this test
Quick answers before you request a test.
Do testers text from real phone numbers?
Yes. Testers use their own phones on real carrier and WhatsApp accounts, so the bot sees ordinary customer numbers and you see ordinary carrier behavior.
Can you test WhatsApp and SMS in the same Audit?
Yes. We can split the panel across channels so you can compare how the same bot behaves on each.
Will testing affect our opt out lists or messaging compliance?
Testers only message numbers and accounts you authorize in writing, and we agree beforehand how test numbers are handled in your opt out records.
Keep reading
Related guides and use cases
More on testing this kind of AI well.
Pre-Launch Testing
A human panel runs your AI through real customer behavior before it goes live.
Read more about Pre-Launch TestingUse caseMultilingual Testing
Native and non-native speakers test whether your AI truly works in every language.
Read more about Multilingual TestingUse caseOngoing Monitoring
A small monthly human panel, a trend line and an alert when your AI slips.
Read more about Ongoing MonitoringGuideHow to Test a Chatbot With Real People
Scope, personas, scenarios, scoring and reading the results, step by step.
Read more about How to Test a Chatbot With Real PeopleGuideThe Chatbot Testing Checklist
A printable list of scenarios and personas for any chatbot.
Read more about The Chatbot Testing ChecklistGuideChatbot KPIs and What They Miss
Containment, deflection, CSAT and the gaps between them.
Read more about Chatbot KPIs and What They MissSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.