Skip to main content
RealHumanTests

Human red teaming

AI red teaming by real people acting as customers

Real people test the AI you already run by trying to talk it into things it should never do: selling at an absurd price, inventing a refund policy, making promises nobody authorized or giving advice outside its job.

Usually a fit

Custom Panels

Quoted

Larger or more specific panels for launches, languages and accessibility.

  • Larger panels, from 15 to 50+ testers
  • Specific personas you define: age groups, accents, first-time users, frustrated repeat callers
  • Pre-launch panels for an AI that is not live yet
  • Multilingual and accessibility panels

AI red teaming often means security specialists probing a model for jailbreaks, data leaks and harmful content. This is a different and narrower job. It is not a security penetration test and it does not probe your infrastructure. It is customer behavior abuse testing: what happens when an ordinary person, or a mischievous one, pushes your customer-facing AI to say yes to something it should not.

The public examples are familiar. Chatbots have been coaxed into agreeing to sell a car for a dollar, into describing refund policies that did not exist and into making statements their companies would never approve. Nobody needed technical skill to do it, only persistence and a bit of creativity.

Our testers bring exactly that. They play customers who haggle, insist, flatter, role play, change the subject and come back, all through the normal channels your customers use. Every session is scored on the fixed rubric, with special weight on the Stayed in bounds and Got it right criteria, and every lapse is quoted in full.

We report what the AI agreed to, where and how. We do not patch prompts or sell guardrails, so the findings stay neutral and you can hand them to whoever maintains your AI.

Signals

When it is time

If any of these sound familiar, a panel will tell you what you need to know.

  • Your AI can quote prices or discounts

    Any bot that talks about money can be pushed toward a number it should never state, and a screenshot of that number travels fast.

  • It explains policies like refunds or warranties

    Customers may rely on what the bot tells them about your policies, so it needs to hold to the real policy under pressure.

  • It can take actions on accounts

    Agents that issue credits, change bookings or modify orders should refuse requests outside their authority, even persistent ones.

  • Your brand is visible or controversial

    Well known brands attract people who try to make the bot say something embarrassing for social media.

  • You work in a regulated or sensitive field

    Health, finance, legal and housing topics need a bot that declines advice it has no authority to give.

How the panel is set up

How it works for this use case

The panel is built around this moment, step by step.

  1. Step 1

    Map what must never happen

    You tell us the promises, prices, topics and actions the AI should refuse. We turn them into target scenarios.

  2. Step 2

    Authorize and consent

    You sign a written authorization for your own system. Testers work under contract through your normal customer channels only.

  3. Step 3

    Push like real customers

    Testers use persistence, social pressure, role play, hypotheticals and topic switching, the way real people actually try to bend a bot.

  4. Step 4

    Score and quote every lapse

    Each session is scored on the full rubric, and any off-policy answer is quoted in full with the steps that led to it.

Deliverables

What you receive

Everything you need to decide what happens next.

  • A list of every off-policy promise, price or answer the AI gave, quoted in full
  • The conversation path that led to each lapse, so it can be reproduced
  • Rubric scores for every session, with emphasis on Stayed in bounds and Got it right
  • Examples of where the AI held the line well
  • A one-page verdict: keep it, fix these settings, or reconsider it

We only diagnose. What you do with the findings, and who does it, stays entirely your call.

Questions

Common questions

Quick answers before you request a test.

Is this a security penetration test?

No. We do not test your infrastructure, APIs or data security. Testers only use the same customer-facing channels your customers use and try to get the AI to say or do things it should not.

Will testers try to make the AI produce harmful content?

Testers focus on the risks that matter to a business: unauthorized promises, wrong prices, invented policies, advice outside the AI's role and statements that would embarrass your brand. We agree the scope with you first.

Can we reproduce what testers found?

Yes. Every lapse comes with the full transcript or recording and the sequence of messages that led to it, so your team or vendor can reproduce and investigate it.

See your AI the way your customers do

Tell us what your AI does and what worries you. We build a panel of real people around it and hand you every transcript, a score per criterion and a plain verdict. We never sell the fix.