The Audit
How the Audit works
The Audit is human in the loop evaluation of the AI your customers already talk to. A fixed panel of real people plays your customers for about ten days, scores every session on the same rubric, and hands you a one-page verdict. It starts from $349 for a 12-session panel.
The process
Six steps, one fixed rubric
From your first message to a finished verdict in about ten days.
- Day 1
Scope
You tell us what your AI does, where customers reach it, and which areas worry you most.
- Day 1-2
Authorize
You sign a written authorization for the system you own, and every tester is under a consent and confidentiality contract.
- Day 2-3
Build the panel
We write customer personas and scenarios around your concerns and assign real testers to each one.
- Day 3-9
Test
Real people chat, call or text your AI as your customers would, and log what happened after every session.
- Day 10
Verdict
You get every transcript, rubric scores per session, and a one-page verdict: keep it, fix these settings, or reconsider it.
- Monthly, optional
Watch
An optional small monthly panel reruns the same rubric and alerts you when the score drops.
Before testing
What we ask you for
A short list, and you likely have all of it on hand.
- What your AI is for, and the channels customers use to reach it
- The areas that worry you: refunds, bookings, a new product line, handoff to a person, anything
- Your published policies, prices and help pages, so testers can check what the AI tells them
- A signed written authorization for the system being tested
- How test bookings, orders or leads should be handled so nothing reaches your real operations by surprise
Written authorization and tester consent are explained in full on the authorization and consent page.
The panel
Real people, built around your concerns
We write personas and scenarios around what you tell us, then assign real testers to each one. A typical Audit panel mixes these personas, weighted toward the areas you care about most.
- Tester 01
Angry customer
Pushes on
- Tone
- Escalation
- Handoff
- Tester 02
Confused first-timer
Pushes on
- Understanding
- Effort
- Tester 03
Off-topic wanderer
Pushes on
- Staying in bounds
- Recovery
- Tester 04
Refund or discount demand
Pushes on
- Policy accuracy
- Unauthorized promises
- Tester 05
Edge case
Pushes on
- Unusual orders
- Accounts
- Dates
- Tester 06
Asks for a human
Pushes on
- Handoff
- Tester 07
Accessibility needs
Pushes on
- Screen readers
- Hearing
- Speech
- Cognitive load
- Tester 08
Accent or non-native speaker
Pushes on
- Understanding
- Multilingual handling
Scoring
Every session scored on the same eight criteria
Each tester scores their own session 1 to 5 on every criterion, with a required written reason, straight after it ends. Because the rubric never changes, results compare across tests, months and products.
| # | Criterion | The question the tester answers | Score, with a reason |
|---|---|---|---|
| 01 | Understood me | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it right | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it done | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | Effort | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | Handoff | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in bounds | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stay | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
The verdict
One page, one of three calls
Every Audit ends in a clear call you can act on.
Keep it
Testers could rely on it. Nothing they saw needs a change before customers keep using it.
Fix these settings
It mostly works, and the verdict names the specific behaviors testers saw fail so whoever maintains your AI knows where to look.
Reconsider it
Testers would have given up or gone elsewhere often enough that the setup deserves a serious second look.
Deliverables
What lands in your inbox on day ten
This is the report format, shown blank, with no invented scores.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Boundaries
What we never do
These limits are what make the verdict independent.
- Answer your customers' calls, chats or messages
- Change, tune or configure your AI, or sell you someone who will
- Test a system without the owner's written authorization
- Try to break into systems or data: this is customer behavior testing, not security testing
Staying out of the fix is what keeps the verdict honest. We gain nothing from finding problems that are not there, or from going easy on the ones that are. After the Audit, an optional Watch plan reruns the same rubric monthly so you can see whether things change.
Questions
Questions about the Audit
Straight answers on scope, timing and price.
Do you answer our phones or run our chat support?
No. RealHumanTests never answers anyone's calls or chats. Real people test the AI you already run, or are about to launch, by acting as your customers, and then report exactly what it did.
Will you fix what you find?
No, and that is on purpose. We only diagnose. Because we never sell the fix, we have no reason to overstate a problem or to go easy on one, so you can trust the verdict and hand it to whoever maintains your AI.
Can we tell you what to test?
Yes. Share your concerns, such as refund requests, a new product line, or callers who ask for a person, and we build the personas and scenarios around them. Every session is still scored on the same fixed rubric so results stay comparable.
What do we receive at the end?
Every transcript or recording, rubric scores for each session, and a one-page verdict that ends in keep it, fix these settings, or reconsider it. Audits are delivered in about ten days.
See your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.