The Benchmark
Voice agent index
A public, human-scored index of AI voice and phone agents. Real callers with real accents, background noise and impatience place the same calls to every product.
Scores
Current status
No score appears here until real people have done the testing.
First edition in testing
No scores are published yet for the Voice Agent Index.
We publish the method, the rubric and the scope first, so anyone can see how scores will be produced before a single one appears. Scores will show here once real people have completed the full scenario set for every product in the first edition. Until then, nothing on this page is a ranking.
Scope
What this index covers
What qualifies for this index, and what does not.
In scope
- AI agents that answer or place customer phone calls for a business
- Voice products sold to businesses, tested on a deployment the owner authorizes
- Publicly reachable business phone lines answered by AI, called only as an ordinary customer would
Out of scope
- Voice assistants on personal devices that are not a business's customer line
- Outbound sales dialers, which customers do not choose to reach
- Any line we cannot call as an ordinary customer or with the owner's written authorization
The scenario set
What every product in this index faces
Every product gets the same scenarios, played by different real people, and every session is scored on the same eight criteria.
- 01A routine request such as hours, an appointment or an order status
- 02A caller with a strong regional or non-native accent
- 03A caller in a noisy place who interrupts and talks over the agent
- 04A caller who stays silent, hesitates or changes their mind mid sentence
- 05A frustrated caller who asks for a person
- 06A request outside what the agent is allowed to promise
Nominations
Suggest a product for this index
Vendors, businesses and anyone else can suggest a product. If you own a deployment and want it included, say so and we will send a written authorization first; nothing is tested without either the owner's authorization or ordinary public customer use, as set out in the Benchmark methodology.
Inclusion is free, and no one can pay for a place, a score or a review before publication. Want this kind of AI tested privately instead? See how we test voice agents for a single company.
See your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.