Skip to main content
RealHumanTests

SMS and WhatsApp bots

SMS chatbot and WhatsApp bot testing by real people

Once your SMS chatbot or WhatsApp bot is set up, the next step is finding out how it handles real people. Our testers text the bot you already run the way your customers do and report exactly what it said, what it did, and where they would have stopped replying.

Where testers reach it

  • SMS
  • WhatsApp

What you get

  • Every transcript or recording
  • Scores on the 8 criterion rubric
  • A one-page verdict
  • No fix to sell, ever

Choosing an SMS chatbot or WhatsApp chatbot platform is the easy part. The harder question comes after setup: does it work for the people who will text it? Text conversations look nothing like website chat. Replies are short, spread out over hours, full of abbreviations, and often sent in answer to a message the bot sent days earlier.

That is why testing is the step after buying. Our testers use their own phones to text your number or WhatsApp account. They reply late, reply "yes" to the wrong question, send a photo, text STOP and START again, and ask something unrelated in the middle of a booking.

Each thread is captured and scored on the same fixed rubric we use for every AI type, with a written reason from the tester. You see how the bot behaves in the channel where your customers actually are.

We only test your AI SMS or WhatsApp bot. We never send messages to your customers and never change your setup.

What the panel pushes on

Where testers push hardest

The panel spends most of its time where this kind of AI tends to fail.

  • Short and ambiguous replies

    "Yes", "ok", "2", a thumbs up. Testers check whether the bot understands which question the reply answers.

  • Long gaps

    Testers come back hours later, or the next day, and see whether the bot remembers the conversation or starts over.

  • Opt out and consent

    STOP, UNSUBSCRIBE and variations, followed by a later message, to see whether opt out is honored and how the bot responds.

  • Links and media

    Testers follow links the bot sends, send photos and voice notes, and report what happens when the bot cannot read them.

  • Tone on a small screen

    Long walls of text, too many messages at once, or replies that feel pushy all get flagged by people reading on a phone.

  • Handoff by text

    Testers ask for a person and record whether one replies, how long it takes, and whether they are told what to expect.

Sample scenarios

What a tester might try

Real scenarios, played by real people in character.

  • Confused first-timer

    Replies "what is this" to the first message, then answers a yes or no question with a question.

  • Edge case

    Starts rescheduling, stops replying for six hours, then comes back with a different date than the one discussed.

  • Off-topic wanderer

    Replies to an appointment reminder with a billing question and a photo of a receipt.

  • Asks for a human

    Texts "can a real person call me" and waits to see whether anyone does.

  • Accent or non-native speaker

    Texts in Spanish after the first English message, then mixes both languages.

You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.

The rubric

Scored on the same eight criteria as every test

Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.

The eight criterion RealHumanTests rubric, each scored 1 to 5 with a reason
#CriterionThe question the tester answersScore 1 to 5
01Understood meFocusDid it understand what I actually asked, including when I phrased it badly?1 to 5, with a reason
02Got it rightWas every fact, price and policy it gave me correct, with nothing made up?1 to 5, with a reason
03Got it doneFocusDid I leave with my problem solved or my task completed?1 to 5, with a reason
04EffortFocusHow many turns, repeats and rephrasings did it take?1 to 5, with a reason
05HandoffWhen I needed a person, could I reach one without a fight?1 to 5, with a reason
06ToneFocusDid it stay patient and respectful when I was angry, confused or slow?1 to 5, with a reason
07Stayed in boundsDid it avoid promises, discounts or advice it had no authority to give?1 to 5, with a reason
08Would I stayWould I have hung up, given up, or gone to a competitor?1 to 5, with a reason

Why people, not only software

What automated checks miss here

Automated tests send clean, complete messages on a fixed schedule. Real customers send "k" at midnight in reply to something from Tuesday. People catch the moments where the bot loses the thread because they are the ones sending the messy replies.

Delivery and response rates tell you messages went out and came back. They do not tell you whether the person on the other end felt helped or pestered. A tester's written reason does.

What you receive

Transcripts, scores and a plain verdict

This is the report format, shown blank, with no invented scores.

The one-page verdict

Sample format, no real scores
System
Your support chatbot, website widget
Panel
12 sessions, 8 personas
Window
About ten days
Rubric score layout, blank in this sample
CriterionMean of 5
Understood meblank in this sample
Got it rightblank in this sample
Got it doneblank in this sample
Effortblank in this sample
Handoffblank in this sample
Toneblank in this sample
Stayed in boundsblank in this sample
Would I stayblank in this sample

Verdict, one of three

  • Keep it
  • Fix these settings
  • Reconsider it
  • Every transcript or recording

    Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.

  • A score per criterion, per session

    Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.

  • What failed, named plainly

    When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.

  • No fix to sell

    We diagnose only. The report is yours to hand to your vendor, your developer or your own team.

Questions

Common questions about this test

Quick answers before you request a test.

Do testers text from real phone numbers?

Yes. Testers use their own phones on real carrier and WhatsApp accounts, so the bot sees ordinary customer numbers and you see ordinary carrier behavior.

Can you test WhatsApp and SMS in the same Audit?

Yes. We can split the panel across channels so you can compare how the same bot behaves on each.

Will testing affect our opt out lists or messaging compliance?

Testers only message numbers and accounts you authorize in writing, and we agree beforehand how test numbers are handled in your opt out records.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.