Guide
Chatbot KPIs, and what they miss without a human review
Containment, deflection, resolution and CSAT are useful numbers. Each can also look great while customers quietly give up.
By RealHumanTests Updated 11 minute read
Chatbot KPIs are the numbers most teams use to judge whether a customer-facing chatbot or AI agent is working: containment, deflection, resolution, escalation, CSAT and a handful of others. They are worth tracking. Every one of them can also look healthy while customers quietly give up. This guide explains each common chatbot KPI, how it is usually calculated, the specific way it can mislead, and what a structured human review tells you that the dashboard cannot.
Definitions vary between platforms, so treat the formulas below as the common form rather than a standard. The first step in reading any dashboard is finding out exactly how your vendor defines each number.
Why chatbot KPIs need a second look
Most chatbot KPIs are computed from what the system can observe: whether a conversation was handed to a person, whether the customer clicked a rating, how long the session lasted. The system cannot observe what the customer did next. It does not see the customer who closed the tab and called your phone line, emailed support, or bought from someone else.
That creates a consistent blind spot. A conversation that ends without a handoff looks like a success to the dashboard whether the customer was helped or simply abandoned the attempt. Keep that in mind as you read each metric below.
Containment rate
Common formula: conversations that ended without a transfer to a human, divided by total conversations.
What it tells you: how much of the conversation volume the bot handled on its own.
How it misleads: containment counts a customer who gave up the same as a customer who was helped. It can also be raised by making human handoff harder to reach, which improves the number while making the experience worse. A bot that refuses to connect people to a person will have excellent containment.
What to pair it with: resolution evidence, repeat contacts, and a sample of contained conversations read by a person.
Deflection rate
Common formula: support contacts avoided because the customer was helped by the bot or self-service content, divided by total support demand. In practice it is often estimated, for example from customers who used the bot and did not open a ticket within a set window.
What it tells you: an estimate of load taken off your human team.
How it misleads: “did not open a ticket” is not the same as “was helped.” Customers who give up entirely, or switch to a different channel the metric does not track, look like deflections. Short measurement windows make this worse.
Resolution rate
Common formula: conversations marked resolved, divided by total conversations. “Resolved” might come from the customer confirming, from the bot deciding, or from the absence of a follow-up.
What it tells you: potentially the most meaningful number on the dashboard, if the definition of resolved is honest.
How it misleads: when the bot itself decides the issue is resolved, the metric inherits the bot’s misunderstandings. If it answered the wrong question confidently, it may also mark that wrong answer as a resolution. Ask who or what sets the resolved flag.
Escalation or handoff rate
Common formula: conversations transferred to a human, divided by total conversations. It is often just the complement of containment.
What it tells you: how often the bot could not or should not finish on its own.
How it misleads: a low escalation rate is usually read as good, but some escalations are exactly right. A billing dispute, a safety concern or a very upset customer should reach a person quickly. The number cannot tell a necessary handoff from an unnecessary one, and it says nothing about how hard the customer had to fight to get it. That is why our rubric scores Handoff directly: when I needed a person, could I reach one without a fight?
Fallback rate
Common formula: bot turns where it fell back to a default response (such as “Sorry, I did not understand”) divided by total bot turns, or conversations with at least one fallback divided by total conversations.
What it tells you: how often the bot knows it did not understand. This was a central metric for older intent-based bots.
How it misleads: generative AI agents rarely say they did not understand. They are more likely to produce a fluent answer to a slightly different question. So a very low fallback rate on a modern agent can hide misunderstanding rather than prove comprehension. The failure has moved from “I do not know” to a confident wrong answer.
Customer: can I change the delivery address on an order that already shipped
Bot: You can update your default delivery address anytime in Account Settings under Addresses.
Tester note: Understood me: 2. Answered how to change my saved address, not a shipped order. No error, no fallback, still wrong.
CSAT and thumbs up or down
Common formula: positive ratings divided by total ratings, usually from a survey or thumbs prompt at the end of the chat.
What it tells you: how the customers who chose to respond felt at that moment.
How it misleads: only some customers answer an in-chat survey, and it is rarely all of them. The customers who abandon mid conversation never see the survey at all. The people most likely to have had a bad experience are the least likely to be counted. CSAT is also collected before the customer knows whether the answer was right: a confident, friendly answer that turns out to be wrong can earn a thumbs up.
Average handle time and conversation length
Common formula: total conversation duration divided by number of conversations. Some dashboards also report average turns per conversation.
What it tells you: how long interactions take.
How it misleads: shorter is not always better. A very short conversation can mean a fast answer or a customer who gave up after two messages. A long one can mean a complex task done well or a customer rephrasing the same question six times. Length needs context before it means anything, which is why the rubric scores Effort from the customer’s side: how many turns, repeats and rephrasings did it take?
Repeat contact rate
Common formula: customers who contact support again about the same issue within a set window after a bot conversation, divided by customers who had a bot conversation.
What it tells you: one of the best available signals of whether the bot actually resolved things, because it looks at what customers did next.
How it misleads: it only works if you can connect the same customer across channels. Anonymous website chats, different phone numbers or email addresses, and customers who simply leave all fall through the gaps. It is worth the effort to measure, but it will undercount.
What a human review adds
Every metric above is an aggregate of signals the system could observe. A structured human review adds the missing piece: a real person inside the conversation, playing a customer with a real goal, who records what happened and what they would have done next.
A good human review is structured, not casual. It has:
- Defined personas such as an angry customer, a confused first-timer, a refund demand, someone who asks for a person, a non-native speaker and someone with accessibility needs.
- Planned scenarios built around the areas you are most worried about.
- A fixed rubric so scores compare across sessions and months. Ours has eight criteria: Understood me, Got it right, Got it done, Effort, Handoff, Tone, Stayed in bounds and Would I stay. See the published rubric.
- A written reason for every score and the full transcript or recording, so anyone can check the judgment.
Here is how the human rubric lines up against the dashboard numbers it complements.
| Dashboard KPI | Blind spot | Rubric criterion that covers it |
|---|---|---|
| Containment, deflection | Cannot tell helped from gave up | Got it done, Would I stay |
| Resolution | Inherits the bot’s own judgment | Got it done, Got it right |
| Escalation | Cannot tell a needed handoff from a failed one | Handoff |
| Fallback | Misses confident wrong answers | Understood me, Got it right |
| CSAT | Misses people who quit before the survey | Tone, Would I stay |
| Handle time | Cannot tell efficient from abandoned | Effort |
| None | Unauthorized promises and discounts | Stayed in bounds |
That last row matters. No standard dashboard metric flags a bot that promised a refund it had no authority to give. People find those, especially when they are deliberately trying to. See human red teaming and our collection of documented chatbot failures.
Human review is also the best way to notice when good numbers start to slip for reasons nobody on your team caused, such as a vendor model update. Our guide on AI drift explains why that happens, and ongoing monitoring describes a small monthly panel that tracks the same rubric over time.
A KPI review checklist
Use this the next time you review your chatbot dashboard.
- Do I know exactly how my vendor defines each metric, especially containment and resolution?
- Who or what sets the resolved flag: the customer, the bot, or the absence of a follow-up?
- Could containment be high because handoff is hard to reach?
- What share of conversations actually answered the CSAT survey?
- Is a low fallback rate hiding confident wrong answers?
- Can I connect bot conversations to repeat contacts in other channels?
- When did someone last read a random sample of contained conversations end to end?
- When did real people last play my customers and score the experience on a fixed rubric?
- Did any of these numbers move after a vendor or model update I did not make?
RealHumanTests supplies the human side of that review. Real people play your customers, score every session on the same rubric and tell you in plain language whether to keep it, fix specific settings or reconsider it. We only diagnose and never sell the fix. For the full method, see how it works, or read how to test a chatbot with real people to run a version yourself.