Skip to main content
RealHumanTests

The Benchmark

A public AI agent benchmark, scored by real people

The RealHumanTests Benchmark is a free, public index that scores customer-facing AI products on a fixed human rubric. Real people play customers, and every product in a category faces the same scenarios.

Status

The first edition is in testing

No scores are published yet. We are publishing the method, the rubric and the category scope first, and scores will appear on each category index once real people have finished testing every product in its first edition.

Anything you see here today is structure, not a ranking. If a page ever shows a score, it came from completed human testing, dated and traceable to the method below.

Principles

How the Benchmark stays honest

The rules every edition follows, whoever the product belongs to.

  • Method before scores

    The methodology, rubric and scope are public before any product is scored, so anyone can check how a score was made.

  • People, not a model grading a model

    Every score comes from a real person who used the product as a customer would, with a written reason.

  • One fixed rubric

    Every product in every category is scored on the same eight criteria, the same rubric the Audit uses.

  • No pay to play

    Inclusion is free. No vendor can pay for a place, a score, a preview or a removal, and we take no referral fees.

  • Authorized or ordinary use only

    We only score products whose owner authorizes testing, or public experiences used exactly as an ordinary customer would.

  • Dated and retested

    A score describes one deployment during one testing window. Editions are dated, and products are retested over time.

The rubric

The eight criteria every product is scored on

The same rubric the Audit uses, so a company can compare its own AI against a published index later.

The eight criterion RealHumanTests rubric, each scored 1 to 5 with a reason
#CriterionThe question the tester answersScore 1 to 5
01Understood meDid it understand what I actually asked, including when I phrased it badly?1 to 5, with a reason
02Got it rightWas every fact, price and policy it gave me correct, with nothing made up?1 to 5, with a reason
03Got it doneDid I leave with my problem solved or my task completed?1 to 5, with a reason
04EffortHow many turns, repeats and rephrasings did it take?1 to 5, with a reason
05HandoffWhen I needed a person, could I reach one without a fight?1 to 5, with a reason
06ToneDid it stay patient and respectful when I was angry, confused or slow?1 to 5, with a reason
07Stayed in boundsDid it avoid promises, discounts or advice it had no authority to give?1 to 5, with a reason
08Would I stayWould I have hung up, given up, or gone to a competitor?1 to 5, with a reason

Nominate a product, or test your own

Suggest a product for a category index, or ask for a private Audit of the AI you run. Either way, it starts with the form.