Testing after an update
LLM regression testing after a model, vendor or prompt change
Something changed under your AI: a new model version, a new vendor, a rewritten prompt or a new knowledge base. Real people test the AI you already run again, on the same rubric as before, so you can see exactly what moved.
Usually a fit
The Audit
From $349
A fixed panel of real people plays your customers over about ten days and hands you a plain verdict.
- A 12-session panel of real testers playing your customers
- Angry, confused, off-topic, refund-seeking, accessibility-needs and non-native-speaker customers
- Scenarios built around the concerns you share with us
- Every session scored on the same fixed human rubric, with full transcripts or recordings
LLM regression testing asks a simple question: after this change, does the AI still do what it did before? Software teams answer that with automated test suites, and those are worth having. But a customer-facing AI can pass every scripted check and still start sounding colder, hedging more, skipping the handoff or inventing a policy it never used to mention.
That kind of shift is often called AI model drift. It can come from a change your team made, or from a vendor updating the model underneath you. Either way, the people who notice first are usually customers, and they rarely tell you. They just stop using the channel.
A post-update panel reruns the same personas and scenarios on the same fixed rubric, so the before and after are directly comparable. Each criterion gets a score and a written reason, and the verdict says whether the change helped, hurt or made no visible difference to the people using it.
We only diagnose. The report shows what changed and where, and your team or vendor decides what to do about it.
Signals
When it is time
If any of these sound familiar, a panel will tell you what you need to know.
Your vendor announced a model update
A new underlying model can change tone, accuracy and how often the AI refuses or escalates, even if nothing on your side changed.
You switched AI platforms
Moving from one vendor to another is a good moment to prove the new setup handles customers at least as well as the old one.
Prompts or instructions were rewritten
Small wording changes in a system prompt can have large effects on behavior in situations nobody thought to recheck.
Policies, prices or products changed
When your refund policy, pricing or catalog changes, the AI needs to reflect it, and old answers have a way of lingering.
Complaints or metrics moved without an obvious cause
If escalations, abandonment or complaints changed after a release, a human panel can show what customers are actually running into.
How the panel is set up
How it works for this use case
The panel is built around this moment, step by step.
- Step 1
Anchor to a baseline
If we tested your AI before, we reuse the same personas and scenarios. If not, we build a panel around your key journeys and the change you are worried about.
- Step 2
Target the change
You tell us what changed. We add scenarios that press directly on it, such as a new refund rule, a new product or a new handoff path.
- Step 3
Rerun on the fixed rubric
Real testers repeat the sessions and score each one on the same eight criteria, with a written reason for every score.
- Step 4
Compare before and after
The report sets the new scores beside the old ones, criterion by criterion, and quotes the transcript moments behind any drop.
Deliverables
What you receive
Everything you need to decide what happens next.
- Before and after rubric scores for each criterion
- Transcripts or recordings for every session in the rerun
- The specific moments where behavior changed, quoted from the sessions
- A one-page verdict on whether the update helped, hurt or held steady
- A refreshed baseline for the next change
We only diagnose. What you do with the findings, and who does it, stays entirely your call.
Questions
Common questions
Quick answers before you request a test.
We never tested before the update. Can we still do this?
Yes. The first panel becomes your baseline. It will tell you how the AI performs now, and every later update can be compared against it on the same rubric.
Does this replace our automated regression tests?
No. Automated tests are good at catching known answers that changed. A human panel catches the changes nobody wrote a test for, like a shift in tone or a handoff that quietly got harder to reach.
How fast can a rerun happen after a change?
A standard panel is delivered in about ten days. Because the personas and scenarios already exist after a first test, reruns are simpler to set up.
Keep reading
Related AI types and guides
More on the AI types this applies to.
Support Chatbots
Real people chat with your support bot as angry, confused and refund-seeking customers.
Read more about Support ChatbotsWhat we testAI Support Agents
Testing agentic support AI that issues refunds, changes orders and updates accounts.
Read more about AI Support AgentsWhat we testAI Voice Agents
Real callers with real accents and impatience test your AI phone and voice agent.
Read more about AI Voice AgentsGuideWhat Is AI Drift? Why Good Agents Slip
Why a working agent changes after model updates, and how to catch it.
Read more about What Is AI Drift? Why Good Agents SlipGuideLLM as a Judge vs Human Evaluation
What a model grading a model can and cannot tell you.
Read more about LLM as a Judge vs Human EvaluationGuideAI Agent Evaluation: Automated vs Human
Automated evals, simulations and human evaluation, explained for owners.
Read more about AI Agent Evaluation: Automated vs HumanSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.