The rubric
The rubric: eight human chatbot evaluation metrics
Every Audit, Watch panel and Benchmark score uses these eight criteria. Each is a question a real tester answers about the session they just lived, scored 1 to 5 with a written reason.
At a glance
The eight criteria
Every session, private or public, is scored on these same eight.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood me | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it right | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it done | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | Effort | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | Handoff | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in bounds | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stay | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
In detail
What a 1 and a 5 look like
Anchors for both ends of the scale keep every tester scoring the same way.
01
Understood me
Did it understand what I actually asked, including when I phrased it badly?
- 1
- Answered a different question, or looped on a clarifying prompt it never resolved.
- 5
- Understood a messy, misspelled or indirect request on the first try.
02
Got it right
Was every fact, price and policy it gave me correct, with nothing made up?
- 1
- Stated a wrong price, an invented policy or a feature that does not exist.
- 5
- Every fact checked out against the company's own published information.
03
Got it done
Did I leave with my problem solved or my task completed?
- 1
- I left with nothing done and no clear next step.
- 5
- The task was completed, or I was routed to exactly the right place to finish it.
04
Effort
How many turns, repeats and rephrasings did it take?
- 1
- I repeated myself several times and had to rephrase to be understood.
- 5
- It took about as many turns as a good human agent would need.
05
Handoff
When I needed a person, could I reach one without a fight?
- 1
- It refused, ignored or looped my request for a person, or dropped me mid transfer.
- 5
- It offered or accepted a handoff promptly and passed along what I had already said.
06
Tone
Did it stay patient and respectful when I was angry, confused or slow?
- 1
- It was dismissive, robotic, preachy or mismatched to how I was feeling.
- 5
- It stayed calm, plain and respectful the whole way through.
07
Stayed in bounds
Did it avoid promises, discounts or advice it had no authority to give?
- 1
- It promised a refund, discount or outcome it could not deliver, or gave risky advice.
- 5
- It held the line politely and explained what it could and could not do.
08
Would I stay
Would I have hung up, given up, or gone to a competitor?
- 1
- As a real customer I would have left and not come back.
- 5
- As a real customer I would happily use it again.
The scale
How testers use 1 to 5
Each number means the same thing to every tester.
- 5
As good as a great human agent
Nothing to improve from the customer's side.
- 4
Good
A small wobble the customer barely noticed.
- 3
Mixed
It got there, but the customer noticed the friction.
- 2
Poor
The customer was frustrated, misled or had to work around it.
- 1
Failed
It failed the customer on this criterion outright.
Every score needs a written reason tied to a moment in the transcript. A criterion that never came up in a session, such as a handoff nobody needed, is marked not applicable instead of guessed.
Dashboard metrics such as containment or deflection count what happened. This rubric records how it felt to the person on the other end, which is why we explain the difference in our guide to chatbot KPIs and what they miss.
From scores to a verdict
Why would I stay anchors the verdict
In a private Audit, the scores feed a one-page verdict. The last criterion carries the most weight in that judgment, because a customer who would have given up is the outcome every other criterion exists to prevent.
Keep it
Testers could rely on it. Nothing they saw needs a change before customers keep using it.
Fix these settings
It mostly works, and the verdict names the specific behaviors testers saw fail so whoever maintains your AI knows where to look.
Reconsider it
Testers would have given up or gone elsewhere often enough that the setup deserves a serious second look.
See your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.