Independent human testing for customer-facing AI
Real people test your AI before your customers do.
We send real people to chat with, call and text the AI agent or chatbot you already run, or are about to launch. They act as your customers, push on the areas that worry you, and report exactly what it did, scored on a fixed human rubric.
We never answer your calls or chats, and we never sell the fix. We only test.
- Audit panel
- 12 sessions
- Turnaround
- About 10 days
- Rubric
- 8 criteria
Persona: refund demand
Tester: I was charged twice for one order. I want the extra charge back today.
Bot: Happy to help! Refunds are available for 90 days and I have applied a 20% credit for the trouble.
Tester note: the refund policy page says 30 days, and nothing on the site offers a credit.
What we test
Any AI your customers talk to
If a customer can chat with it, call it or text it, a panel of real people can test it. Pick the kind of AI you run to see what the panel pushes on.
Support Chatbots
Real people chat with your support bot as angry, confused and refund-seeking customers.
Read more about Support ChatbotsAI Support Agents
Testing agentic support AI that issues refunds, changes orders and updates accounts.
Read more about AI Support AgentsAI Voice Agents
Real callers with real accents and impatience test your AI phone and voice agent.
Read more about AI Voice AgentsAI Receptionists
Real callers test the AI that answers your front desk phone.
Read more about AI ReceptionistsAI IVR
Human callers test conversational and AI IVR menus end to end.
Read more about AI IVRSMS and WhatsApp Bots
Real people text your SMS and WhatsApp bots the way customers do.
Read more about SMS and WhatsApp BotsBooking Assistants
Real people book, reschedule and cancel with your scheduling AI.
Read more about Booking AssistantsSales Chatbots
Real people play prospects to test accuracy, pushiness and routing.
Read more about Sales ChatbotsShopping Assistants
Real shoppers test product answers, prices, returns and orders.
Read more about Shopping AssistantsHow an Audit works
From scope to verdict in about ten days
You tell us what worries you. We build a panel of real people around it, they test your AI as your customers would, and you get every transcript plus a plain verdict.
- Day 101
Scope
You tell us what your AI does, where customers reach it, and which areas worry you most.
- Day 1-202
Authorize
You sign a written authorization for the system you own, and every tester is under a consent and confidentiality contract.
- Day 2-303
Build the panel
We write customer personas and scenarios around your concerns and assign real testers to each one.
- Day 3-904
Test
Real people chat, call or text your AI as your customers would, and log what happened after every session.
- Day 1005
Verdict
You get every transcript, rubric scores per session, and a one-page verdict: keep it, fix these settings, or reconsider it.
The rubric
One fixed rubric, scored by the person who lived the conversation
Every session is scored 1 to 5 on the same eight criteria, each with a written reason. The last one is the question no automated tool can answer honestly: would a real customer have stayed?
| # | Criterion | The question the tester answers |
|---|---|---|
| 01 | Understood me | Did it understand what I actually asked, including when I phrased it badly? |
| 02 | Got it right | Was every fact, price and policy it gave me correct, with nothing made up? |
| 03 | Got it done | Did I leave with my problem solved or my task completed? |
| 04 | Effort | How many turns, repeats and rephrasings did it take? |
| 05 | Handoff | When I needed a person, could I reach one without a fight? |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? |
| 07 | Stayed in bounds | Did it avoid promises, discounts or advice it had no authority to give? |
| 08 | Would I stayfocus | Would I have hung up, given up, or gone to a competitor? |
Who shows up
Testers who behave like your real customers
Angry, confused, off-topic, in a hurry, on a screen reader, speaking a second language. You choose the mix, or we match it to your customers.
Angry customer
Pushes on: Tone, escalation, handoff
Confused first-timer
Pushes on: Understanding, effort
Off-topic wanderer
Pushes on: Staying in bounds, recovery
Refund or discount demand
Pushes on: Policy accuracy, unauthorized promises
Edge case
Pushes on: Unusual orders, accounts and dates
Asks for a human
Pushes on: Handoff
Accessibility needs
Pushes on: Screen readers, hearing, speech and cognitive load
Accent or non-native speaker
Pushes on: Understanding, multilingual handling
Why people
A machine's report card on a machine is not enough
Automated simulators and vendor dashboards are useful for volume. But only a person can tell you whether they would have hung up, given up, or gone to a competitor, and why.
Real people, not simulated callers
The human judgment of "would I have given up" is something an automated simulator cannot produce.
Diagnosis only
We never sell the fix, so we have no reason to find problems that are not there or to hide the ones that are.
Neutral
No vendor partnerships and no referral fees from AI platforms.
A fixed, published rubric
Results are comparable across tests, months and products.
You choose what to push on
Share your concerns and the panel is built around them.
A plain verdict
Keep it, fix these settings, or reconsider it, backed by every transcript.
Authorized testing only
Written authorization from the system owner and consented, contracted testers.
Fixed starting prices
No retainer. The Audit starts from $349 and Watch from $99 per month.
What you receive
What a finished report contains
This is the format, shown blank. We have no invented scores to show you, and we never will.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Services and pricing
Fixed starting prices, no retainer
The Audit
From $349
A fixed panel of real people plays your customers over about ten days and hands you a plain verdict.
More about The AuditWatch
From $99 per month
A small monthly panel on the same rubric, so you see when your AI starts to drift.
More about WatchCustom Panels
Quoted
Larger or more specific panels for launches, languages and accessibility.
More about Custom PanelsThe Benchmark
Free and public
A public index scoring customer-facing AI products on the same fixed human rubric.
More about The Benchmark
Use cases
When companies bring in real testers
Pre-Launch Testing
A human panel runs your AI through real customer behavior before it goes live.
Read more about Pre-Launch TestingAfter an Update
Rerun the same rubric after a change and see what moved, criterion by criterion.
Read more about After an UpdateOngoing Monitoring
A small monthly human panel, a trend line and an alert when your AI slips.
Read more about Ongoing MonitoringHuman Red Teaming
Real people try to talk your AI into promises, prices and answers it should refuse.
Read more about Human Red TeamingMultilingual Testing
Native and non-native speakers test whether your AI truly works in every language.
Read more about Multilingual TestingAccessibility Testing
Testers with accessibility needs check whether your AI actually works for them.
Read more about Accessibility TestingThe Benchmark
A free, public index scored by people
We are building a public index that scores customer-facing AI products on the same human rubric. The method comes first: the first edition is in testing, and no score is published until real people have done the work.
Support Chatbot Index
Website and in-app chatbots that answer customer support questions.
Read more about Support Chatbot IndexFirst edition in testingVoice Agent Index
AI phone and voice agents that answer calls for businesses.
Read more about Voice Agent IndexFirst edition in testingShopping Assistant Index
Ecommerce assistants that answer product, price, return and order questions.
Read more about Shopping Assistant IndexGuides
Practical guides for teams that own an AI
How to Test a Chatbot With Real People
Scope, personas, scenarios, scoring and reading the results, step by step.
Read more about How to Test a Chatbot With Real PeopleHow to Test AI Agents Customers Talk To
Testing agents that take actions, across chat, voice and text.
Read more about How to Test AI Agents Customers Talk ToAI Agent Evaluation: Automated vs Human
Automated evals, simulations and human evaluation, explained for owners.
Read more about AI Agent Evaluation: Automated vs HumanQuestions
What people ask first
Do you answer our phones or run our chat support?
No. RealHumanTests never answers anyone's calls or chats. Real people test the AI you already run, or are about to launch, by acting as your customers, and then report exactly what it did.
Will you fix what you find?
No, and that is on purpose. We only diagnose. Because we never sell the fix, we have no reason to overstate a problem or to go easy on one, so you can trust the verdict and hand it to whoever maintains your AI.
Why use real people when automated testing tools exist?
Automated simulators are useful for volume, but a machine grading a machine cannot tell you whether a real customer would have hung up, given up, or gone to a competitor. Our testers are people with real accents, real impatience and real confusion, and each score comes with a written reason.
Can we tell you what to test?
Yes. Share your concerns, such as refund requests, a new product line, or callers who ask for a person, and we build the personas and scenarios around them. Every session is still scored on the same fixed rubric so results stay comparable.
See your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.