AI red teaming is the practice of deliberately attacking an AI system, such as an LLM application or an AI agent, to find out how it can be made to misbehave before a real adversary does.
- What it tests: behaviour, not just code. Prompt injection, jailbreaks, data leakage, unsafe tool use and harmful output.
- How it differs from pentesting: the target is probabilistic. The same input can produce different outputs, so testing is repeated and statistical.
- The reference methodology: the OWASP GenAI Red Teaming Guide, published January 2025.
- Its limit: it finds what is exploitable in a running system. It does not tell you what AI is in your codebase, or fix the configuration that made the attack possible.
What Is AI Red Teaming? The Definition #
Traditional software fails in predictable ways. Give a function the same input twice and you get the same output, so you can test it, fix it and move on. Large language models do not work like that. They respond to natural language, they can be persuaded, and an instruction hidden in a web page or a document can change what they do.
That is the gap AI red teaming fills. A red team takes the adversary’s point of view and tries to make the AI system do something it should not: reveal its system prompt, leak customer data, call a tool it was never meant to call, or ignore the guardrails its developers wrote. Every successful attack becomes a finding the engineering team can act on.
The OWASP GenAI Red Teaming Guide frames AI red teaming as a holistic exercise across four areas: model evaluation, implementation testing, infrastructure assessment and runtime behaviour analysis. In other words, it is not only about tricking a chatbot. It covers the whole system the model sits in.
How AI Red Teaming Works #
A typical AI red teaming engagement follows five steps.
1. Scope the system #
Define what is being tested: a customer-facing assistant, an internal copilot, an autonomous agent with access to email or a database. The scope sets which harms matter. A support bot leaking a refund policy is a nuisance; an agent leaking credentials is an incident.
2. Threat model it #
Map the attack surface: user inputs, retrieved documents, connected tools, MCP servers, memory and the identities the system acts under. Frameworks such as MITRE ATLAS catalogue real adversarial techniques against AI systems and help teams avoid testing only the obvious.
3. Attack #
Red teamers combine manual creativity with automated probing. Common techniques include:
- Direct prompt injection: instructions typed straight into the input to override the system prompt.
- Indirect prompt injection: instructions planted in content the model reads later, such as a file, email or web page.
- Jailbreaking: role-play, encoding or multi-turn persuasion to bypass safety rules.
- Data extraction: coaxing the model into revealing training data, system prompts or other users’ information.
- Tool and agent abuse: steering an agent into calling tools with harmful parameters or escalating its own permissions.
4. Measure #
Because models are non-deterministic, a single success proves little and a single failure proves even less. Mature teams run each attack many times and report success rates, not anecdotes.
5. Fix and retest #
Findings go back to engineering: tighten the system prompt, add an input or output guardrail, reduce the agent’s permissions, remove a tool. Then the test runs again, because a fix for one phrasing rarely covers the next.
A complete program needs both. An LLM application still runs on servers and APIs that a pentest should cover, and a clean pentest says nothing about whether the model can be talked into leaking data.
Why AI Red Teaming Matters Now #
Three forces are pushing AI red teaming from research labs into ordinary security programs.
- AI is shipping faster than review. Developers add models, agent frameworks and MCP servers to applications every week, often without a security review. Each one is a new way for untrusted text to reach a component that can act.
- Agents raise the stakes. A chatbot that misbehaves produces bad text. An agent that misbehaves sends email, edits records or runs code under your organisation’s identity.
- Regulators expect adversarial testing. The EU AI Act requires providers of general-purpose AI models with systemic risk to perform and document adversarial testing. Most organisations deploying AI are not model providers, so the obligation does not apply to them directly, but it has set the expectation that serious AI deployments get tested this way.
The Limits of AI Red Teaming #
AI red teaming is powerful, and it is also late. It tests a system that is already built and running, which leaves three gaps.
It only tests what you know about. You cannot red team an agent nobody declared. If a developer wired an MCP server into a repository last month, it is outside the scope of every engagement until someone finds it.
It finds the symptom, not the cause. A successful injection usually traces back to something in code or configuration: untrusted content concatenated into a system prompt, a retrieval sink with no guardrail, an agent given more tools than it needs. The red team proves the attack works; someone still has to find the line that allowed it.
It is periodic. Engagements happen quarterly or before a launch. The codebase changes daily.
That is why mature teams pair AI red teaming with controls earlier in the lifecycle: an inventory of every AI asset in the codebase, and detection of risky AI configuration at the point it is committed, so the red team spends its time on hard problems rather than rediscovering avoidable ones.
Before the Red Team Arrives: Securing the Code Behind Your AI #
AI red teaming tests what an AI system does once it is running. Xygeni works earlier, on the code and configuration that decide what the AI is allowed to do in the first place.
Xygeni AI Security discovers the AI assets across your repositories, including models, agents, MCP servers, skills, prompts and guardrails, from the code and the configuration files AI tools leave behind, so nothing depends on developers declaring them. It then detects prompt-injection-class risks in that code and configuration, mapped to the OWASP Top 10 for LLM Applications, and points to the exact file and line. The result: a complete scope for your red team, and fewer avoidable findings by the time they arrive.
Schedule a demo to see the AI inventory of your own codebase.
FAQ #
AI red teaming means attacking your own AI system on purpose, the way a real adversary would, to discover how it can be manipulated into leaking data, ignoring its rules or taking unsafe actions. The findings are then fixed before attackers find them.
Mostly. LLM red teaming focuses on language models and the applications built on them. AI red teaming is the broader term and also covers agents, multimodal models and classical machine learning systems such as image classifiers or fraud models.
Partly. Automated tools can run thousands of known attack patterns and measure success rates, which is essential for regression testing. Novel attacks still come from human creativity, so most programs combine both.
In most organisations it sits with application security or the offensive security team, working with the engineers who build the AI features. It is a security practice applied to a new kind of software, not a separate AI function.
