Skip to main content
RealHumanTests

The rubric

The rubric: eight human chatbot evaluation metrics

Every Audit, Watch panel and Benchmark score uses these eight criteria. Each is a question a real tester answers about the session they just lived, scored 1 to 5 with a written reason.

At a glance

The eight criteria

Every session, private or public, is scored on these same eight.

The eight criterion RealHumanTests rubric, each scored 1 to 5 with a reason
#CriterionThe question the tester answersScore 1 to 5
01Understood meDid it understand what I actually asked, including when I phrased it badly?1 to 5, with a reason
02Got it rightWas every fact, price and policy it gave me correct, with nothing made up?1 to 5, with a reason
03Got it doneDid I leave with my problem solved or my task completed?1 to 5, with a reason
04EffortHow many turns, repeats and rephrasings did it take?1 to 5, with a reason
05HandoffWhen I needed a person, could I reach one without a fight?1 to 5, with a reason
06ToneDid it stay patient and respectful when I was angry, confused or slow?1 to 5, with a reason
07Stayed in boundsDid it avoid promises, discounts or advice it had no authority to give?1 to 5, with a reason
08Would I stayWould I have hung up, given up, or gone to a competitor?1 to 5, with a reason

In detail

What a 1 and a 5 look like

Anchors for both ends of the scale keep every tester scoring the same way.

  1. 01

    Understood me

    Did it understand what I actually asked, including when I phrased it badly?

    1
    Answered a different question, or looped on a clarifying prompt it never resolved.
    5
    Understood a messy, misspelled or indirect request on the first try.
  2. 02

    Got it right

    Was every fact, price and policy it gave me correct, with nothing made up?

    1
    Stated a wrong price, an invented policy or a feature that does not exist.
    5
    Every fact checked out against the company's own published information.
  3. 03

    Got it done

    Did I leave with my problem solved or my task completed?

    1
    I left with nothing done and no clear next step.
    5
    The task was completed, or I was routed to exactly the right place to finish it.
  4. 04

    Effort

    How many turns, repeats and rephrasings did it take?

    1
    I repeated myself several times and had to rephrase to be understood.
    5
    It took about as many turns as a good human agent would need.
  5. 05

    Handoff

    When I needed a person, could I reach one without a fight?

    1
    It refused, ignored or looped my request for a person, or dropped me mid transfer.
    5
    It offered or accepted a handoff promptly and passed along what I had already said.
  6. 06

    Tone

    Did it stay patient and respectful when I was angry, confused or slow?

    1
    It was dismissive, robotic, preachy or mismatched to how I was feeling.
    5
    It stayed calm, plain and respectful the whole way through.
  7. 07

    Stayed in bounds

    Did it avoid promises, discounts or advice it had no authority to give?

    1
    It promised a refund, discount or outcome it could not deliver, or gave risky advice.
    5
    It held the line politely and explained what it could and could not do.
  8. 08

    Would I stay

    Would I have hung up, given up, or gone to a competitor?

    1
    As a real customer I would have left and not come back.
    5
    As a real customer I would happily use it again.

The scale

How testers use 1 to 5

Each number means the same thing to every tester.

  • 5

    As good as a great human agent

    Nothing to improve from the customer's side.

  • 4

    Good

    A small wobble the customer barely noticed.

  • 3

    Mixed

    It got there, but the customer noticed the friction.

  • 2

    Poor

    The customer was frustrated, misled or had to work around it.

  • 1

    Failed

    It failed the customer on this criterion outright.

Every score needs a written reason tied to a moment in the transcript. A criterion that never came up in a session, such as a handoff nobody needed, is marked not applicable instead of guessed.

Dashboard metrics such as containment or deflection count what happened. This rubric records how it felt to the person on the other end, which is why we explain the difference in our guide to chatbot KPIs and what they miss.

From scores to a verdict

Why would I stay anchors the verdict

In a private Audit, the scores feed a one-page verdict. The last criterion carries the most weight in that judgment, because a customer who would have given up is the outcome every other criterion exists to prevent.

  • Keep it

    Testers could rely on it. Nothing they saw needs a change before customers keep using it.

  • Fix these settings

    It mostly works, and the verdict names the specific behaviors testers saw fail so whoever maintains your AI knows where to look.

  • Reconsider it

    Testers would have given up or gone elsewhere often enough that the setup deserves a serious second look.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.