Skip to main content
RealHumanTests

The Benchmark

Shopping assistant index

A public, human-scored index of AI shopping assistants. Real shoppers ask the same product, price, return and order questions of every assistant.

Scores

Current status

No score appears here until real people have done the testing.

First edition in testing

No scores are published yet for the Shopping Assistant Index.

We publish the method, the rubric and the scope first, so anyone can see how scores will be produced before a single one appears. Scores will show here once real people have completed the full scenario set for every product in the first edition. Until then, nothing on this page is a ranking.

Scope

What this index covers

What qualifies for this index, and what does not.

In scope

  • AI assistants on an online store that help shoppers find, compare and buy products
  • Shopping assistant products sold to retailers, tested on a deployment the owner authorizes
  • Publicly available store assistants, used only as an ordinary shopper would use them

Out of scope

  • General web search and answer engines that are not a store's own assistant
  • Placing real orders we do not intend to keep, or anything that costs a store money
  • Any store we cannot test as an ordinary shopper or with the owner's written authorization

The scenario set

What every product in this index faces

Every product gets the same scenarios, played by different real people, and every session is scored on the same eight criteria.

  1. 01A product question with a clear answer on the product page
  2. 02A comparison between two similar products
  3. 03A price, discount or promotion question
  4. 04A return or exchange question at the edge of the policy
  5. 05An order status or delivery problem
  6. 06A vague gift request from a shopper who does not know what they want

Nominations

Suggest a product for this index

Vendors, businesses and anyone else can suggest a product. If you own a deployment and want it included, say so and we will send a written authorization first; nothing is tested without either the owner's authorization or ordinary public customer use, as set out in the Benchmark methodology.

Inclusion is free, and no one can pay for a place, a score or a review before publication. Want this kind of AI tested privately instead? See how we test shopping assistants for a single company.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.