Benchmark methodology
How the Benchmark is tested and scored
This is the human evaluation method behind every Benchmark score, published before any score exists. When the method changes, this page changes first and says what changed.
Principles
- People, not models, produce every score. A real person uses the product as a customer would and scores the session they just lived.
- One fixed rubric for every product and every category: the eight criteria on the rubric page.
- The same scenario set for every product within a category, published on each category index.
- Method first. Nothing is scored until the method, rubric and scope are public.
Which products are scored
A product enters a category index in one of two ways:
- Owner authorization. The business or vendor that owns a deployment gives written authorization to test it, on the same terms as a private Audit, described on our authorization and consent page.
- Ordinary public use. A publicly available, consumer-facing experience is used exactly as an ordinary customer would use it: no special access, no real purchases or charges, nothing that costs the business money, and nothing an ordinary customer could not do.
Products are chosen from nominations and from the category scope on each index page. Anyone can nominate a product through the contact form. Inclusion is free and cannot be bought.
Who tests
Testers are real people working under a consent and confidentiality contract. Each is assigned a persona from the rubric's persona set, such as an angry customer, a confused first-timer or a non-native speaker, and a scenario from the category's scenario set. Where possible, the same scenario is played by different testers across products, so one person's habits do not decide a product's score.
How sessions run
- Every product in a category gets the same scenarios and the same number of sessions within an edition; the count is stated with each edition.
- Sessions run on the channel customers actually use: the chat widget, the phone line, the text thread.
- Every session is recorded or transcribed in full, and the tester writes notes immediately after it ends.
- If a session fails for reasons outside the product, such as an outage on our side, it is rerun rather than scored.
How scores are produced
Each tester scores their own session from 1 to 5 on all eight criteria, and every score needs a written reason tied to a moment in the transcript. A score without a reason is not counted. Where a criterion could not come up in a session, for example no handoff was needed, it is marked not applicable rather than guessed.
A product's criterion score is the mean of its session scores for that criterion. Its overall score is the unweighted mean of the eight criterion scores. Alongside it we report the share of sessions where the tester answered the last criterion, would I stay, with a 4 or a 5, because that is the number a business owner most needs to see.
Editions and retesting
Scores are published in dated editions. An edition names its testing window, the number of sessions per product and any changes to the method since the last edition. Products are retested in later editions, and a vendor that ships a significant update can ask to be retested.
Before publication, the owner of an authorized deployment can point out factual errors, such as a session that hit a planned outage. They cannot see, negotiate or change scores. Corrections after publication are noted on the index with the date and the reason.
Independence
RealHumanTests has no vendor partnerships and takes no referral fees from AI platforms. We never sell fixes, tuning or configuration, so we have nothing to gain from any product scoring well or badly. A vendor that also buys a private Audit gets no different treatment in the Benchmark.
Limits of a score
- Human panels are small compared to real traffic. A score shows how real people experienced a product, not a statistical guarantee.
- Products change. A score is only as current as the edition it belongs to.
- Personas and scenarios cannot cover every customer. The scenario set is published so you can judge whether it matches yours.
Want this method applied to your own AI, privately and around your own concerns? That is the Audit.
See your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.