The Benchmark
A public AI agent benchmark, scored by real people
The RealHumanTests Benchmark is a free, public index that scores customer-facing AI products on a fixed human rubric. Real people play customers, and every product in a category faces the same scenarios.
Status
The first edition is in testing
No scores are published yet. We are publishing the method, the rubric and the category scope first, and scores will appear on each category index once real people have finished testing every product in its first edition.
Anything you see here today is structure, not a ranking. If a page ever shows a score, it came from completed human testing, dated and traceable to the method below.
Category indexes
Three categories to start, more over time
Each index publishes its scope and scenario set before any score.
Support Chatbot Index
Website and in-app chatbots that answer customer support questions.
Read more about Support Chatbot IndexFirst edition in testingVoice Agent Index
AI phone and voice agents that answer calls for businesses.
Read more about Voice Agent IndexFirst edition in testingShopping Assistant Index
Ecommerce assistants that answer product, price, return and order questions.
Read more about Shopping Assistant IndexPrinciples
How the Benchmark stays honest
The rules every edition follows, whoever the product belongs to.
Method before scores
The methodology, rubric and scope are public before any product is scored, so anyone can check how a score was made.
People, not a model grading a model
Every score comes from a real person who used the product as a customer would, with a written reason.
One fixed rubric
Every product in every category is scored on the same eight criteria, the same rubric the Audit uses.
No pay to play
Inclusion is free. No vendor can pay for a place, a score, a preview or a removal, and we take no referral fees.
Authorized or ordinary use only
We only score products whose owner authorizes testing, or public experiences used exactly as an ordinary customer would.
Dated and retested
A score describes one deployment during one testing window. Editions are dated, and products are retested over time.
The rubric
The eight criteria every product is scored on
The same rubric the Audit uses, so a company can compare its own AI against a published index later.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood me | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it right | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it done | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | Effort | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | Handoff | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in bounds | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stay | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
Nominate a product, or test your own
Suggest a product for a category index, or ask for a private Audit of the AI you run. Either way, it starts with the form.