Skip to main content
RealHumanTests

Ongoing monitoring: the Watch plan

Chatbot monitoring by real people, every month

Watch sends a small panel of real people to test the AI you already run every month. They play your customers on the same fixed rubric, so you get a trend line per criterion and an alert when the score drops.

Usually a fit

Watch

From $99 per month

A small monthly panel on the same rubric, so you see when your AI starts to drift.

  • Scores comparable month to month on the same rubric
  • A trend line per rubric criterion
  • An alert when the score drops after a vendor model update or prompt change

Most chatbot monitoring and AI agent monitoring tools watch the system: uptime, latency, error rates, token costs and conversation counts. LLM monitoring platforms go further and flag unusual outputs. All of that is useful, and none of it tells you whether a real customer on a real day would have given up.

Customer-facing AI changes over time even when your team changes nothing. Vendors ship new model versions, knowledge bases get edited and prompts get tuned for one problem at the expense of another. The result can be an agent that looked great at launch and slowly becomes harder to use.

Watch is a small panel of real testers who revisit your AI each month with the same personas and scenarios, scored on the same fixed rubric. Because nothing about the test moves, any change in the scores reflects a change in the AI. You see a trend line for each criterion, and we alert you when the score drops.

Watch starts from $99 per month. It diagnoses only: the alert tells you what testers saw change, and your team or vendor decides what to do next.

Signals

When it is time

If any of these sound familiar, a panel will tell you what you need to know.

  • Your AI is live and customers depend on it

    Once real customers rely on a chatbot or voice agent, a slow decline costs more than a sudden outage because nobody notices it.

  • Your vendor updates the model without much notice

    Hosted AI platforms change their underlying models over time. A monthly human check shows whether those changes reached your customers.

  • Several people edit prompts or content

    When marketing, support and engineering all touch the AI's instructions or knowledge base, a stable outside check keeps everyone honest.

  • You report AI performance to leadership

    A monthly human score on a published rubric is easier to explain than a dashboard of containment rates.

  • You completed an Audit and want to keep the baseline

    Watch reuses the Audit's rubric, so the first month's trend starts from a known point.

How the panel is set up

How it works for this use case

The panel is built around this moment, step by step.

  1. Step 1

    Set the fixed panel

    We agree on a small set of personas and scenarios that cover your most important customer journeys. They stay the same every month.

  2. Step 2

    Test every month

    Real testers run the sessions and score each one on the eight rubric criteria, each with a written reason.

  3. Step 3

    Plot the trend

    Scores go on a trend line per criterion, so you can see whether understanding, accuracy, handoff or tone is moving.

  4. Step 4

    Alert on a drop

    When a score falls, you get an alert that names the criterion and quotes the sessions behind it.

Deliverables

What you receive

Everything you need to decide what happens next.

  • Monthly rubric scores from the same fixed panel design
  • A trend line for each of the eight criteria
  • An alert when the score drops, with the sessions that caused it
  • Transcripts or recordings from every monthly session
  • A short monthly note on anything testers noticed that changed

We only diagnose. What you do with the findings, and who does it, stays entirely your call.

Questions

Common questions

Quick answers before you request a test.

How is Watch different from our analytics dashboard?

Your dashboard counts what happened across all conversations. Watch shows what a real person experienced in a controlled set of scenarios, scored the same way every month, so a change in the score means the AI changed.

Can we change the scenarios later?

Yes, when your products or policies change. We keep the core scenarios stable so the trend line stays meaningful, and add new ones alongside them.

Do we need an Audit before Watch?

It helps, because the Audit sets a detailed baseline, but it is not required. The first month of Watch can serve as the starting point.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.