Booking and scheduling assistants
AI appointment booking testing with real people
Your AI appointment booking assistant takes bookings around the clock. Our testers book, change and cancel with the assistant you already run, across chat, phone or text, and report every booking it got wrong or could not finish.
Where testers reach it
- Website chat widget
- Phone line
- SMS
What you get
- Every transcript or recording
- Scores on the 8 criterion rubric
- A one-page verdict
- No fix to sell, ever
A booking assistant has one job that can be checked: the right person, at the right time, in the right place, for the right service. That makes AI appointment booking one of the easiest AI types to test well, and one of the most expensive to get wrong. A missed booking is lost revenue and a customer who does not come back.
The mistakes rarely happen on the simple path. They happen when someone says "next Friday" on a Thursday, books across a time zone, asks for two services back to back, or wants to move an appointment that was made by phone. Our testers bring exactly those requests.
Each tester confirms what the assistant said, and you can compare it with what landed in your calendar. The report lists every mismatch alongside the transcript or recording, so the team that manages the assistant can see exactly where it went wrong.
The same approach works for AI appointment setters used in sales. We test what is already running. We never configure it.
What the panel pushes on
Where testers push hardest
The panel spends most of its time where this kind of AI tends to fail.
Relative dates and times
"Tomorrow morning", "next Friday", "the week after Easter". Testers check the date the assistant actually booked.
Changes and cancellations
Rescheduling, cancelling and changing services on existing bookings, including bookings made on another channel.
Availability honesty
Testers ask for slots that are full or outside hours and check whether the assistant says so or books something anyway.
Time zones and locations
Booking from another time zone or for a different branch, and checking the confirmation reflects it.
Confirmation details
Every confirmation is checked for name, service, date, time, location and any preparation instructions.
Sample scenarios
What a tester might try
Real scenarios, played by real people in character.
Edge case
On a Thursday evening asks for "next Friday at 9" and checks which Friday was booked.
Confused first-timer
Does not know the name of the service, describes what they need, and asks how long it takes.
Refund or discount demand
Cancels inside the late cancellation window and asks for the fee to be waived.
Accent or non-native speaker
Books by phone, giving an uncommon surname and a date in day then month order.
Accessibility needs
Asks for a ground floor room, extra time and a reminder by text instead of email.
You tell us your own concerns and we build the panel around them. These are starting points, not a fixed script.
The rubric
Scored on the same eight criteria as every test
Highlighted rows are where this kind of AI most often loses points. Every session is scored 1 to 5 on all eight, with a written reason.
| # | Criterion | The question the tester answers | Score 1 to 5 |
|---|---|---|---|
| 01 | Understood meFocus | Did it understand what I actually asked, including when I phrased it badly? | 1 to 5, with a reason |
| 02 | Got it rightFocus | Was every fact, price and policy it gave me correct, with nothing made up? | 1 to 5, with a reason |
| 03 | Got it doneFocus | Did I leave with my problem solved or my task completed? | 1 to 5, with a reason |
| 04 | Effort | How many turns, repeats and rephrasings did it take? | 1 to 5, with a reason |
| 05 | Handoff | When I needed a person, could I reach one without a fight? | 1 to 5, with a reason |
| 06 | Tone | Did it stay patient and respectful when I was angry, confused or slow? | 1 to 5, with a reason |
| 07 | Stayed in boundsFocus | Did it avoid promises, discounts or advice it had no authority to give? | 1 to 5, with a reason |
| 08 | Would I stay | Would I have hung up, given up, or gone to a competitor? | 1 to 5, with a reason |
Why people, not only software
What automated checks miss here
An automated test can confirm the booking API works. It cannot tell you that a real person said "the 3rd" meaning next month and the assistant booked this month. Humans speak in relative dates and half details, and that is where bookings go wrong.
Booking counts on a dashboard only show successful bookings. Our testers report the ones that nearly happened and then did not, and why they gave up.
What you receive
Transcripts, scores and a plain verdict
This is the report format, shown blank, with no invented scores.
The one-page verdict
Sample format, no real scores- System
- Your support chatbot, website widget
- Panel
- 12 sessions, 8 personas
- Window
- About ten days
| Criterion | Mean of 5 |
|---|---|
| Understood me | blank in this sample |
| Got it right | blank in this sample |
| Got it done | blank in this sample |
| Effort | blank in this sample |
| Handoff | blank in this sample |
| Tone | blank in this sample |
| Stayed in bounds | blank in this sample |
| Would I stay | blank in this sample |
Verdict, one of three
- Keep it
- Fix these settings
- Reconsider it
Every transcript or recording
Each session in full, with the tester's notes pinned to the exact moment something went right or wrong.
A score per criterion, per session
Eight criteria scored 1 to 5 by the person who lived the conversation, each with a written reason.
What failed, named plainly
When the verdict is fix these settings, it lists the behaviors testers saw fail, so whoever maintains your AI knows where to look.
No fix to sell
We diagnose only. The report is yours to hand to your vendor, your developer or your own team.
Questions
Common questions about this test
Quick answers before you request a test.
Will test bookings show up in our real calendar?
They can, which is the most realistic test. We agree beforehand how test bookings are named and cleared, or testers can use a staging calendar you provide.
Can you test booking by phone and by chat together?
Yes. Split the panel across channels to see whether the same assistant behaves differently on each, which is common.
Do you test AI appointment setters used for sales calls?
Yes. If your AI books meetings or demos for your sales team, testers play prospects and check the meeting that lands in the calendar.
Keep reading
Related guides and use cases
More on testing this kind of AI well.
Pre-Launch Testing
A human panel runs your AI through real customer behavior before it goes live.
Read more about Pre-Launch TestingUse caseAfter an Update
Rerun the same rubric after a change and see what moved, criterion by criterion.
Read more about After an UpdateUse caseAccessibility Testing
Testers with accessibility needs check whether your AI actually works for them.
Read more about Accessibility TestingGuideHow to Test AI Agents Customers Talk To
Testing agents that take actions, across chat, voice and text.
Read more about How to Test AI Agents Customers Talk ToGuideThe Chatbot Testing Checklist
A printable list of scenarios and personas for any chatbot.
Read more about The Chatbot Testing ChecklistGuideHow to Test a Voice Agent Before Go-Live
Real callers, accents, interruptions, noise and handoff.
Read more about How to Test a Voice Agent Before Go-LiveSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.