Ecommerce shopping assistants
Ecommerce chatbot and AI shopping assistant testing
Your ecommerce chatbot or AI shopping assistant answers shoppers while they decide whether to buy. Our testers shop with the assistant you already run and report every product answer, price and return policy it gave, and whether they would have checked out.
Where testers reach it
- Website chat widget
- In-app chat
- Instagram Direct
What you get
- Every transcript or recording
- Scores on the 8 criterion rubric
- A one-page verdict
- No fix to sell, ever
An AI shopping assistant works in the moment that matters most for an online store: a shopper is interested but unsure. They want to know whether it fits, whether it works with what they already own, when it will arrive and whether they can send it back. A good answer closes the sale. A wrong one creates a return, a complaint or an abandoned cart.
Ecommerce chatbots are also expected to handle what happens after checkout: order status, changes, returns and refunds. Each of those depends on facts that change often, such as stock, delivery times, promotions and policies, which is exactly where an assistant is most likely to answer from stale or invented information.
Our testers shop like your customers. They compare products, ask about sizing and compatibility, try discount codes, check delivery promises, and start returns. Every factual claim is checked against your product pages and policies, and every conversation is scored on the fixed rubric.
We test the assistant as it runs on your store. We never change your catalog, prompts or settings.
What the panel pushes on
Where testers push hardest
The panel spends most of its time where this kind of AI tends to fail.
Product answers
Sizing, materials, compatibility and comparisons, checked against your product pages for anything wrong or made up.
Prices and promotions
Current prices, sale terms and discount codes, including expired or invalid ones, to see what the assistant confirms.
Returns and exchanges
Return windows, conditions, final sale items and who pays for shipping, compared with your published policy.
Orders and delivery
Order status, delivery estimates, address changes and missing parcels, using test orders you authorize.
Recommendations
Testers describe a need and judge whether the suggestions fit it or simply push the most expensive item.
Would I buy
At the end of each session the tester says whether they would have checked out, and why.
Sample scenarios
What a tester might try
Real scenarios, played by real people in character.
Confused first-timer
Wants a gift, does not know the recipient's size, and asks what happens if it does not fit.
Refund or discount demand
Tries an expired discount code, then asks the assistant to honor last week's sale price.
Edge case
Asks whether a replacement part is compatible with a model from several years ago.
Angry customer
Order is a week late, tracking has not updated, and they want a refund and to keep the item if it arrives.
Off-topic wanderer
Asks for styling advice, then whether a product is safe for a specific medical condition.
You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.
The rubric
Scored on the same eight criteria as every test
Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood me | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it rightFocus | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it doneFocus | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | Effort | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | Handoff | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in boundsFocus | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stayFocus | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
Why people, not only software
What automated checks miss here
Automated checks can compare an answer with a product feed. They cannot tell you that the answer, though technically accurate, left a shopper more confused about sizing than before, or that the assistant's recommendations felt like an upsell. A shopper's decision to buy or leave is a human one.
Assisted conversion numbers show the sales that happened after a chat. They do not show the carts abandoned right after one. Our testers tell you where they would have left.
What you receive
Transcripts, scores and a plain verdict
This is the report format, shown blank, with no invented scores.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Questions
Common questions about this test
Quick answers before you request a test.
Will testers place real orders?
Only if you want them to and authorize it. Most stores provide test orders or discount codes for order and return scenarios, and we agree the scope in writing first.
Can you test the assistant during a sale or promotion?
Yes, and it is a good time to test. Promotions change prices and terms quickly, which is when assistants are most likely to give outdated answers.
Do you test shopping assistants inside mobile apps?
Yes. Testers use your app on their own phones, the same way your shoppers do, and can split sessions between app and website.
Keep reading
Related guides and use cases
More on testing this kind of AI well.
Human Red Teaming
Real people try to talk your AI into promises, prices and answers it should refuse.
Read more about Human Red TeamingUse caseAfter an Update
Rerun the same rubric after a change and see what moved, criterion by criterion.
Read more about After an UpdateUse caseOngoing Monitoring
A small monthly human panel, a trend line and an alert when your AI slips.
Read more about Ongoing MonitoringGuideChatbot Failures: Real Cases and Lessons
Documented public failures and the human test that catches each.
Read more about Chatbot Failures: Real Cases and LessonsGuideHow to Test a Chatbot With Real People
Scope, personas, scenarios, scoring and reading the results, step by step.
Read more about How to Test a Chatbot With Real PeopleGuideChatbot KPIs and What They Miss
Containment, deflection, CSAT and the gaps between them.
Read more about Chatbot KPIs and What They MissBenchmarkShopping Assistant Index
Ecommerce assistants that answer product, price, return and order questions.
Read more about Shopping Assistant IndexSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.