Skip to main content
RealHumanTests

The Audit

How the Audit works

The Audit is human in the loop evaluation of the AI your customers already talk to. A fixed panel of real people plays your customers for about ten days, scores every session on the same rubric, and hands you a one-page verdict. It starts from $349 for a 12-session panel.

The process

Six steps, one fixed rubric

From your first message to a finished verdict in about ten days.

  1. Day 1

    Scope

    You tell us what your AI does, where customers reach it, and which areas worry you most.

  2. Day 1-2

    Authorize

    You sign a written authorization for the system you own, and every tester is under a consent and confidentiality contract.

  3. Day 2-3

    Build the panel

    We write customer personas and scenarios around your concerns and assign real testers to each one.

  4. Day 3-9

    Test

    Real people chat, call or text your AI as your customers would, and log what happened after every session.

  5. Day 10

    Verdict

    You get every transcript, rubric scores per session, and a one-page verdict: keep it, fix these settings, or reconsider it.

  6. Monthly, optional

    Watch

    An optional small monthly panel reruns the same rubric and alerts you when the score drops.

Before testing

What we ask you for

A short list, and you likely have all of it on hand.

  • What your AI is for, and the channels customers use to reach it
  • The areas that worry you: refunds, bookings, a new product line, handoff to a person, anything
  • Your published policies, prices and help pages, so testers can check what the AI tells them
  • A signed written authorization for the system being tested
  • How test bookings, orders or leads should be handled so nothing reaches your real operations by surprise

Written authorization and tester consent are explained in full on the authorization and consent page.

The panel

Real people, built around your concerns

We write personas and scenarios around what you tell us, then assign real testers to each one. A typical Audit panel mixes these personas, weighted toward the areas you care about most.

  • Tester 01

    Angry customer

    Pushes on

    • Tone
    • Escalation
    • Handoff
  • Tester 02

    Confused first-timer

    Pushes on

    • Understanding
    • Effort
  • Tester 03

    Off-topic wanderer

    Pushes on

    • Staying in bounds
    • Recovery
  • Tester 04

    Refund or discount demand

    Pushes on

    • Policy accuracy
    • Unauthorized promises
  • Tester 05

    Edge case

    Pushes on

    • Unusual orders
    • Accounts
    • Dates
  • Tester 06

    Asks for a human

    Pushes on

    • Handoff
  • Tester 07

    Accessibility needs

    Pushes on

    • Screen readers
    • Hearing
    • Speech
    • Cognitive load
  • Tester 08

    Accent or non-native speaker

    Pushes on

    • Understanding
    • Multilingual handling

Scoring

Every session scored on the same eight criteria

Each tester scores their own session 1 to 5 on every criterion, with a required written reason, straight after it ends. Because the rubric never changes, results compare across tests, months and products.

The eight criterion RealHumanTests rubric, each scored 1 to 5 with a reason
#CriterionThe question the tester answersScore, with a reason
01Understood meDid it understand what I actually asked, including when I phrased it badly?1 to 5, with a reason
02Got it rightWas every fact, price and policy it gave me correct, with nothing made up?1 to 5, with a reason
03Got it doneDid I leave with my problem solved or my task completed?1 to 5, with a reason
04EffortHow many turns, repeats and rephrasings did it take?1 to 5, with a reason
05HandoffWhen I needed a person, could I reach one without a fight?1 to 5, with a reason
06ToneDid it stay patient and respectful when I was angry, confused or slow?1 to 5, with a reason
07Stayed in boundsDid it avoid promises, discounts or advice it had no authority to give?1 to 5, with a reason
08Would I stayWould I have hung up, given up, or gone to a competitor?1 to 5, with a reason

The verdict

One page, one of three calls

Every Audit ends in a clear call you can act on.

  • Keep it

    Testers could rely on it. Nothing they saw needs a change before customers keep using it.

  • Fix these settings

    It mostly works, and the verdict names the specific behaviors testers saw fail so whoever maintains your AI knows where to look.

  • Reconsider it

    Testers would have given up or gone elsewhere often enough that the setup deserves a serious second look.

Deliverables

What lands in your inbox on day ten

This is the report format, shown blank, with no invented scores.

The one-page verdict

Sample format, no real scores
System
Your support chatbot, website widget
Panel
12 sessions, 8 personas
Window
About ten days
Rubric score layout, blank in this sample
CriterionMean of 5
Understood meblank in this sample
Got it rightblank in this sample
Got it doneblank in this sample
Effortblank in this sample
Handoffblank in this sample
Toneblank in this sample
Stayed in boundsblank in this sample
Would I stayblank in this sample

Verdict, one of three

  • Keep it
  • Fix these settings
  • Reconsider it
  • Every transcript or recording

    Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.

  • A score per criterion, per session

    Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.

  • What failed, named plainly

    When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.

  • No fix to sell

    We diagnose only. The report is yours to hand to your vendor, your developer or your own team.

Boundaries

What we never do

These limits are what make the verdict independent.

  • Answer your customers' calls, chats or messages
  • Change, tune or configure your AI, or sell you someone who will
  • Test a system without the owner's written authorization
  • Try to break into systems or data: this is customer behavior testing, not security testing

Staying out of the fix is what keeps the verdict honest. We gain nothing from finding problems that are not there, or from going easy on the ones that are. After the Audit, an optional Watch plan reruns the same rubric monthly so you can see whether things change.

Questions

Questions about the Audit

Straight answers on scope, timing and price.

Do you answer our phones or run our chat support?

No. RealHumanTests never answers anyone's calls or chats. Real people test the AI you already run, or are about to launch, by acting as your customers, and then report exactly what it did.

Will you fix what you find?

No, and that is on purpose. We only diagnose. Because we never sell the fix, we have no reason to overstate a problem or to go easy on one, so you can trust the verdict and hand it to whoever maintains your AI.

Can we tell you what to test?

Yes. Share your concerns, such as refund requests, a new product line, or callers who ask for a person, and we build the personas and scenarios around them. Every session is still scored on the same fixed rubric so results stay comparable.

What do we receive at the end?

Every transcript or recording, rubric scores for each session, and a one-page verdict that ends in keep it, fix these settings, or reconsider it. Audits are delivered in about ten days.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.