Skip to main content
RealHumanTests

Testing after an update

LLM regression testing after a model, vendor or prompt change

Something changed under your AI: a new model version, a new vendor, a rewritten prompt or a new knowledge base. Real people test the AI you already run again, on the same rubric as before, so you can see exactly what moved.

Usually a fit

The Audit

From $349

A fixed panel of real people plays your customers over about ten days and hands you a plain verdict.

  • A 12-session panel of real testers playing your customers
  • Angry, confused, off-topic, refund-seeking, accessibility-needs and non-native-speaker customers
  • Scenarios built around the concerns you share with us
  • Every session scored on the same fixed human rubric, with full transcripts or recordings

LLM regression testing asks a simple question: after this change, does the AI still do what it did before? Software teams answer that with automated test suites, and those are worth having. But a customer-facing AI can pass every scripted check and still start sounding colder, hedging more, skipping the handoff or inventing a policy it never used to mention.

That kind of shift is often called AI model drift. It can come from a change your team made, or from a vendor updating the model underneath you. Either way, the people who notice first are usually customers, and they rarely tell you. They just stop using the channel.

A post-update panel reruns the same personas and scenarios on the same fixed rubric, so the before and after are directly comparable. Each criterion gets a score and a written reason, and the verdict says whether the change helped, hurt or made no visible difference to the people using it.

We only diagnose. The report shows what changed and where, and your team or vendor decides what to do about it.

Signals

When it is time

If any of these sound familiar, a panel will tell you what you need to know.

  • Your vendor announced a model update

    A new underlying model can change tone, accuracy and how often the AI refuses or escalates, even if nothing on your side changed.

  • You switched AI platforms

    Moving from one vendor to another is a good moment to prove the new setup handles customers at least as well as the old one.

  • Prompts or instructions were rewritten

    Small wording changes in a system prompt can have large effects on behavior in situations nobody thought to recheck.

  • Policies, prices or products changed

    When your refund policy, pricing or catalog changes, the AI needs to reflect it, and old answers have a way of lingering.

  • Complaints or metrics moved without an obvious cause

    If escalations, abandonment or complaints changed after a release, a human panel can show what customers are actually running into.

How the panel is set up

How it works for this use case

The panel is built around this moment, step by step.

  1. Step 1

    Anchor to a baseline

    If we tested your AI before, we reuse the same personas and scenarios. If not, we build a panel around your key journeys and the change you are worried about.

  2. Step 2

    Target the change

    You tell us what changed. We add scenarios that press directly on it, such as a new refund rule, a new product or a new handoff path.

  3. Step 3

    Rerun on the fixed rubric

    Real testers repeat the sessions and score each one on the same eight criteria, with a written reason for every score.

  4. Step 4

    Compare before and after

    The report sets the new scores beside the old ones, criterion by criterion, and quotes the transcript moments behind any drop.

Deliverables

What you receive

Everything you need to decide what happens next.

  • Before and after rubric scores for each criterion
  • Transcripts or recordings for every session in the rerun
  • The specific moments where behavior changed, quoted from the sessions
  • A one-page verdict on whether the update helped, hurt or held steady
  • A refreshed baseline for the next change

We only diagnose. What you do with the findings, and who does it, stays entirely your call.

Questions

Common questions

Quick answers before you request a test.

We never tested before the update. Can we still do this?

Yes. The first panel becomes your baseline. It will tell you how the AI performs now, and every later update can be compared against it on the same rubric.

Does this replace our automated regression tests?

No. Automated tests are good at catching known answers that changed. A human panel catches the changes nobody wrote a test for, like a shift in tone or a handoff that quietly got harder to reach.

How fast can a rerun happen after a change?

A standard panel is delivered in about ten days. Because the personas and scenarios already exist after a first test, reruns are simpler to set up.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.