Ongoing monitoring: the Watch plan
Chatbot monitoring by real people, every month
Watch sends a small panel of real people to test the AI you already run every month. They play your customers on the same fixed rubric, so you get a trend line per criterion and an alert when the score drops.
Usually a fit
Watch
From $99 per month
A small monthly panel on the same rubric, so you see when your AI starts to drift.
- Scores comparable month to month on the same rubric
- A trend line per rubric criterion
- An alert when the score drops after a vendor model update or prompt change
Most chatbot monitoring and AI agent monitoring tools watch the system: uptime, latency, error rates, token costs and conversation counts. LLM monitoring platforms go further and flag unusual outputs. All of that is useful, and none of it tells you whether a real customer on a real day would have given up.
Customer-facing AI changes over time even when your team changes nothing. Vendors ship new model versions, knowledge bases get edited and prompts get tuned for one problem at the expense of another. The result can be an agent that looked great at launch and slowly becomes harder to use.
Watch is a small panel of real testers who revisit your AI each month with the same personas and scenarios, scored on the same fixed rubric. Because nothing about the test moves, any change in the scores reflects a change in the AI. You see a trend line for each criterion, and we alert you when the score drops.
Watch starts from $99 per month. It diagnoses only: the alert tells you what testers saw change, and your team or vendor decides what to do next.
Signals
When it is time
If any of these sound familiar, a panel will tell you what you need to know.
Your AI is live and customers depend on it
Once real customers rely on a chatbot or voice agent, a slow decline costs more than a sudden outage because nobody notices it.
Your vendor updates the model without much notice
Hosted AI platforms change their underlying models over time. A monthly human check shows whether those changes reached your customers.
Several people edit prompts or content
When marketing, support and engineering all touch the AI's instructions or knowledge base, a stable outside check keeps everyone honest.
You report AI performance to leadership
A monthly human score on a published rubric is easier to explain than a dashboard of containment rates.
You completed an Audit and want to keep the baseline
Watch reuses the Audit's rubric, so the first month's trend starts from a known point.
How the panel is set up
How it works for this use case
The panel is built around this moment, step by step.
- Step 1
Set the fixed panel
We agree on a small set of personas and scenarios that cover your most important customer journeys. They stay the same every month.
- Step 2
Test every month
Real testers run the sessions and score each one on the eight rubric criteria, each with a written reason.
- Step 3
Plot the trend
Scores go on a trend line per criterion, so you can see whether understanding, accuracy, handoff or tone is moving.
- Step 4
Alert on a drop
When a score falls, you get an alert that names the criterion and quotes the sessions behind it.
Deliverables
What you receive
Everything you need to decide what happens next.
- Monthly rubric scores from the same fixed panel design
- A trend line for each of the eight criteria
- An alert when the score drops, with the sessions that caused it
- Transcripts or recordings from every monthly session
- A short monthly note on anything testers noticed that changed
We only diagnose. What you do with the findings, and who does it, stays entirely your call.
Questions
Common questions
Quick answers before you request a test.
How is Watch different from our analytics dashboard?
Your dashboard counts what happened across all conversations. Watch shows what a real person experienced in a controlled set of scenarios, scored the same way every month, so a change in the score means the AI changed.
Can we change the scenarios later?
Yes, when your products or policies change. We keep the core scenarios stable so the trend line stays meaningful, and add new ones alongside them.
Do we need an Audit before Watch?
It helps, because the Audit sets a detailed baseline, but it is not required. The first month of Watch can serve as the starting point.
Keep reading
Related AI types and guides
More on the AI types this applies to.
Support Chatbots
Real people chat with your support bot as angry, confused and refund-seeking customers.
Read more about Support ChatbotsWhat we testAI Support Agents
Testing agentic support AI that issues refunds, changes orders and updates accounts.
Read more about AI Support AgentsWhat we testAI Voice Agents
Real callers with real accents and impatience test your AI phone and voice agent.
Read more about AI Voice AgentsWhat we testSMS and WhatsApp Bots
Real people text your SMS and WhatsApp bots the way customers do.
Read more about SMS and WhatsApp BotsGuideWhat Is AI Drift? Why Good Agents Slip
Why a working agent changes after model updates, and how to catch it.
Read more about What Is AI Drift? Why Good Agents SlipGuideChatbot KPIs and What They Miss
Containment, deflection, CSAT and the gaps between them.
Read more about Chatbot KPIs and What They MissGuideAI Agent Evaluation: Automated vs Human
Automated evals, simulations and human evaluation, explained for owners.
Read more about AI Agent Evaluation: Automated vs HumanSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.