Human red teaming
AI red teaming by real people acting as customers
Real people test the AI you already run by trying to talk it into things it should never do: selling at an absurd price, inventing a refund policy, making promises nobody authorized or giving advice outside its job.
Usually a fit
Custom Panels
Quoted
Larger or more specific panels for launches, languages and accessibility.
- Larger panels, from 15 to 50+ testers
- Specific personas you define: age groups, accents, first-time users, frustrated repeat callers
- Pre-launch panels for an AI that is not live yet
- Multilingual and accessibility panels
AI red teaming often means security specialists probing a model for jailbreaks, data leaks and harmful content. This is a different and narrower job. It is not a security penetration test and it does not probe your infrastructure. It is customer behavior abuse testing: what happens when an ordinary person, or a mischievous one, pushes your customer-facing AI to say yes to something it should not.
The public examples are familiar. Chatbots have been coaxed into agreeing to sell a car for a dollar, into describing refund policies that did not exist and into making statements their companies would never approve. Nobody needed technical skill to do it, only persistence and a bit of creativity.
Our testers bring exactly that. They play customers who haggle, insist, flatter, role play, change the subject and come back, all through the normal channels your customers use. Every session is scored on the fixed rubric, with special weight on the Stayed in bounds and Got it right criteria, and every lapse is quoted in full.
We report what the AI agreed to, where and how. We do not patch prompts or sell guardrails, so the findings stay neutral and you can hand them to whoever maintains your AI.
Signals
When it is time
If any of these sound familiar, a panel will tell you what you need to know.
Your AI can quote prices or discounts
Any bot that talks about money can be pushed toward a number it should never state, and a screenshot of that number travels fast.
It explains policies like refunds or warranties
Customers may rely on what the bot tells them about your policies, so it needs to hold to the real policy under pressure.
It can take actions on accounts
Agents that issue credits, change bookings or modify orders should refuse requests outside their authority, even persistent ones.
Your brand is visible or controversial
Well known brands attract people who try to make the bot say something embarrassing for social media.
You work in a regulated or sensitive field
Health, finance, legal and housing topics need a bot that declines advice it has no authority to give.
How the panel is set up
How it works for this use case
The panel is built around this moment, step by step.
- Step 1
Map what must never happen
You tell us the promises, prices, topics and actions the AI should refuse. We turn them into target scenarios.
- Step 2
Authorize and consent
You sign a written authorization for your own system. Testers work under contract through your normal customer channels only.
- Step 3
Push like real customers
Testers use persistence, social pressure, role play, hypotheticals and topic switching, the way real people actually try to bend a bot.
- Step 4
Score and quote every lapse
Each session is scored on the full rubric, and any off-policy answer is quoted in full with the steps that led to it.
Deliverables
What you receive
Everything you need to decide what happens next.
- A list of every off-policy promise, price or answer the AI gave, quoted in full
- The conversation path that led to each lapse, so it can be reproduced
- Rubric scores for every session, with emphasis on Stayed in bounds and Got it right
- Examples of where the AI held the line well
- A one-page verdict: keep it, fix these settings, or reconsider it
We only diagnose. What you do with the findings, and who does it, stays entirely your call.
Questions
Common questions
Quick answers before you request a test.
Is this a security penetration test?
No. We do not test your infrastructure, APIs or data security. Testers only use the same customer-facing channels your customers use and try to get the AI to say or do things it should not.
Will testers try to make the AI produce harmful content?
Testers focus on the risks that matter to a business: unauthorized promises, wrong prices, invented policies, advice outside the AI's role and statements that would embarrass your brand. We agree the scope with you first.
Can we reproduce what testers found?
Yes. Every lapse comes with the full transcript or recording and the sequence of messages that led to it, so your team or vendor can reproduce and investigate it.
Keep reading
Related AI types and guides
More on the AI types this applies to.
Sales Chatbots
Real people play prospects to test accuracy, pushiness and routing.
Read more about Sales ChatbotsWhat we testShopping Assistants
Real shoppers test product answers, prices, returns and orders.
Read more about Shopping AssistantsWhat we testAI Support Agents
Testing agentic support AI that issues refunds, changes orders and updates accounts.
Read more about AI Support AgentsWhat we testSupport Chatbots
Real people chat with your support bot as angry, confused and refund-seeking customers.
Read more about Support ChatbotsGuideChatbot Failures: Real Cases and Lessons
Documented public failures and the human test that catches each.
Read more about Chatbot Failures: Real Cases and LessonsGuideHow to Test AI Agents Customers Talk To
Testing agents that take actions, across chat, voice and text.
Read more about How to Test AI Agents Customers Talk ToGuideThe Chatbot Testing Checklist
A printable list of scenarios and personas for any chatbot.
Read more about The Chatbot Testing ChecklistSee your AI the way your customers do
Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.