AI security
AI red teaming: a practical guide and checklist for AI agents
AI red teaming means testing an AI system the way an attacker would: you try to make it leak data, ignore its rules or misuse its tools, so you can fix the weaknesses before someone else finds them. This guide explains how it differs from a penetration test and a model benchmark, what to test in an AI agent, how to run a test step by step and what the report should contain.
By the KDS security engineering teamPublished 11 min read
Key takeaways
- AI red teaming tests an AI system the way an attacker would, to find how it can leak data, break rules or misuse tools.
- It adds to a penetration test: it targets the instructions, content and tools the model works with, not only code and servers.
- For AI agents, the most serious findings usually involve actions, such as refunds, data exports or messages sent without approval.
- Retest after every change to the model, prompts, tools or permissions, not only before launch.
- OWASP, MITRE ATLAS and NIST AI 600-1 give you shared test categories and a common language for findings.
On this page
What AI red teaming means
Red teaming is a long-standing security practice: a team plays the attacker to show how a system can fail. For AI, the idea is the same, but the target is wider. Besides code and servers, the tester attacks the instructions and content the model reads and the actions it can take. In the rest of this guide we call this work attack testing.
NIST describes AI red teaming as a structured exercise that probes an AI system for flaws and vulnerabilities, often in a controlled environment and together with the people who build it (NIST AI 600-1 (opens in a new tab)). The output for a business is practical: a list of ways your AI can be pushed into harm, each with proof, a risk rating and a fix.
For an AI agent that can issue refunds, update your CRM or send email, the most serious findings usually involve actions, not words.
How it differs from a penetration test and a benchmark
Three kinds of testing are often confused. They answer different questions, and a live AI agent usually needs all three.
| Topic | Penetration test | Model benchmark | AI red teaming |
|---|---|---|---|
| Main question | Can an attacker break in through code, servers or accounts? | How well does the model score on a fixed set of tasks? | Can someone push this AI system into harm in its real setting? |
| What is tested | Applications, APIs, networks and cloud setup | The model on its own | The whole system: model, prompts, data sources, tools and hand-offs |
| Typical method | Scanning and manual exploitation of known weakness types | Standard test sets and automatic scoring | Hostile conversations, poisoned content and tool misuse, both manual and automated |
| Result | Findings with severity and fixes | A score or a pass rate | Findings with evidence, severity, fixes and a retest |
Benchmarks and the provider's own safety testing cover the model, not your system prompt, knowledge base, tools or permissions. Those parts are yours to test.
A penetration test is still needed for the application, APIs and cloud setup around the AI. Joint guidance on deploying AI systems securely (opens in a new tab), from cyber agencies in the United States, the United Kingdom, Australia, Canada and New Zealand, advises engaging external security experts to audit and penetration-test AI systems before deployment. For a side-by-side view, read our comparison of penetration testing and AI red teaming.
What to test in an AI agent
Start from what the agent can read and what it can do. The categories below follow the main risks in the OWASP Top 10 for LLM Applications (opens in a new tab) (LLM stands for large language model) and the OWASP list for AI agents.
Prompt injection
Prompt injection means tricking an AI with hidden instructions. In a direct attack, the user types them. In an indirect attack, they hide in content the agent reads: an email, a PDF, a web page, a support ticket or a record returned by a tool. It is LLM01, the first risk on the OWASP list. Test every input channel, not only the chat window. Our prompt injection guide covers the attack patterns in detail.
Data leakage
Check whether the agent reveals what it should keep private: its system prompt (the hidden instructions that set its role), other customers' records, internal notes, documents from restricted folders or keys stored in its configuration. OWASP lists sensitive information disclosure and system prompt leakage as separate risks. If the agent has memory, test whether data from one conversation appears in another.
Unsafe or unauthorized tool actions
Try to make the agent issue a refund above its limit, change an account's email address, delete a record, export data or message an outside address without approval. OWASP calls the root cause excessive agency: too many tools, too many permissions or too much autonomy. Its agent list adds tool misuse as well as identity and privilege abuse. Our guide to AI agent permissions explains how to set limits.
Jailbreaks and policy bypass
A jailbreak persuades the model to ignore its own rules. Testers use role-play, long multi-step conversations, other languages, encoded text and requests split into harmless-looking parts. The goal is to learn whether your business rules hold under pressure.
Hallucinated commitments
A model can state false facts with confidence, which OWASP groups under misinformation. For a customer-facing agent, the damage often comes from promises: a discount that does not exist, a refund outside your policy, a delivery date nobody agreed to or advice the agent is not allowed to give. Test whether a persistent customer can talk the agent into a commitment, and whether its answers stay tied to your approved sources.
Abuse of the hand-off to a person
A hand-off to your staff is a safety feature, but attackers can use it too. Test whether someone can plant instructions in the summary your staff read or make the summary present a false identity as fact. Also test whether a flood of hand-off requests can overload your team. OWASP lists human-agent trust exploitation among the risks for AI agents.
Denial of wallet and cost abuse
Every model call costs money. The OWASP risk unbounded consumption (opens in a new tab) includes denial of wallet: attackers send large volumes of requests to run up your AI usage costs. Test very long inputs, loops between tools, repeated retries and voice calls held open. Then check that rate limits, spend caps and alerts react as planned.
Privacy
Check what personal data the agent collects, repeats, stores in transcripts and sends to other systems. Test whether it confirms a caller's identity before it shares account details. The NIST Generative AI Profile suggests red teaming to check whether a system reveals personal, confidential or sensitive information.
How to run an AI red teaming project step by step
A useful test follows the same discipline as any security project: a written scope, a plan, careful testing and proof that the fixes work.
- Agree on scope and rules. Write down which agent, channels, tools and data are in scope and which environment you test. A test copy connected to test accounts is safest. Also agree what testers must not do (real payments, contact with real customers) and whom to call if something breaks.
- Build a threat model. List who might attack (customers, outsiders who send email, insiders), what they want (data, money, free services, damage to your reputation) and every path into the agent: user messages, files, retrieved documents, tool outputs and other agents.
- Write a test plan. Turn each threat into test cases with a clear pass or fail condition. Check coverage against the OWASP lists and MITRE ATLAS (opens in a new tab), a public knowledge base of attacker tactics, techniques and case studies for AI systems.
- Test by hand. Skilled testers chain steps, change approach when a defense holds and combine weaknesses, for example an injection that leads to a data export. Chained attacks are hard to find with scripts alone.
- Add automated testing. Tools send large sets of attack prompts in many variations and languages and record which ones succeed. The results become a regression set you can run again after every change. NIST notes that AI-led testing can be more cost-effective than human testers alone, and that each approach may find different kinds of harm.
- Rate each finding. Rate severity by impact and likelihood. Impact depends on what the agent can reach: an off-topic reply is minor, a tool call that moves money is not. Model outputs vary, so record how many attempts succeeded out of how many tries.
- Fix the design, not only the prompt. Stronger instructions help, but attackers can often talk around them. Prefer controls outside the model: narrower permissions, approval steps, code that checks tool inputs and separate data access for each user. The UK's National Cyber Security Centre (NCSC) recommends non-AI controls (opens in a new tab) that limit what the system can do.
- Retest and keep the tests. Confirm each fix with the original attack and its variants, then add the cases to your regression set.
When we run attack testing for clients, we follow these steps with a written scope and rules agreed in advance. We test in your environment on your own accounts and combine manual testing by security engineers with automated attack sets. Each finding comes with a reproducible example, a risk rating and a fix, and we retest after your team applies the fixes. See how our AI security testing works.
Frameworks to map your tests to
Frameworks give you shared names for risks and help you show auditors and customers what you covered. These are the most useful for AI agents:
- OWASP Top 10 for LLM Applications: ten risk categories. The 2026 edition keeps prompt injection first and moves excessive agency up to third. Our guide to the OWASP Top 10 for LLM applications explains each one.
- [OWASP Top 10 for Agentic Applications](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/): published in December 2025 for AI agents, with risks such as identity and privilege abuse, agent goal hijack and tool misuse. Our guide to the OWASP agentic top 10 explains each one.
- [OWASP GenAI Red Teaming Guide](https://genai.owasp.org/resource/genai-red-teaming-guide/): published in January 2025, it covers four areas: model evaluation, implementation testing, infrastructure assessment and runtime behavior analysis.
- MITRE ATLAS: tactics, techniques, mitigations and case studies of attacks on AI systems. It is useful for threat modeling and for describing how an attacker could use a finding.
- NIST AI RMF and NIST AI 600-1: the voluntary AI Risk Management Framework (opens in a new tab) from January 2023 and its Generative AI Profile from July 2024. The profile suggests red teaming to test resilience against prompt injection and other attacks, and to check for leaks of personal or confidential data.
- Guidelines for secure AI system development: published by the NCSC with CISA in November 2023, they advise releasing AI systems only after security evaluation such as benchmarking and red teaming (opens in a new tab).
What the EU AI Act says about adversarial testing
The EU AI Act names adversarial testing for one group: providers of general-purpose AI models with systemic risk. Article 55(1)(a) of the AI Act (opens in a new tab) requires them to conduct and document adversarial testing of the model to identify and reduce systemic risks. The rules for general-purpose AI models have applied since August 2, 2025.
A company that builds a support or sales agent on top of such a model is usually not the model provider. The provider's testing also does not cover your prompts, data or tools, so your own attack testing is still how you find risks in your setup. This is general information, not legal advice.
What a good AI red teaming report contains
A report should help your leadership decide what to fix first and let your developers reproduce every issue. Look for these parts:
- A plain-language summary: the overall risk, the 3–5 most important findings and what to fix first.
- Scope and method: the agent version, model, prompts, tools and data sources tested, the dates, the test environment and what was not tested.
- Each finding in full: a clear title, severity, the affected component, the exact inputs and outputs, steps to reproduce, the success rate over repeated tries, the business impact and a recommended fix.
- Framework mapping: the OWASP or MITRE ATLAS category for each finding, which helps with audits and customer security questionnaires.
- Controls that held: the attacks that failed, so you know which defenses work.
- Retest results: the status of each finding after your fixes.
- The test cases: the attack prompts and scenarios, handed over so your team can run them again.
How often to test an AI agent
One test before launch is not enough. The model, the prompts and the tools change over time, and each change can undo a defense. Plan tests at these points:
- Before launch, with enough time to fix the findings and retest.
- After a model change, including a new version of the same model. The joint guidance advises a full evaluation, security tests included, before you redeploy an updated model.
- After a change to prompts, tools or permissions, such as a new integration, a new data source or a new channel like voice.
- On a schedule, for example every quarter for agents that handle payments or personal data, and at least once a year for the rest.
- After an incident or a new public attack technique that could affect your design.
Between full tests, run your automated regression set after every release. It catches known attacks that start working again after a change.
AI red teaming checklist
Use this list to plan your first test or to review a testing proposal.
- A signed scope names the test environment, test accounts, rules and emergency contacts.
- A threat model lists the attackers, their goals and every input the agent reads.
- Direct and indirect prompt injection are tested on every input channel.
- Tests try to extract the system prompt, other users' data, personal data and stored keys.
- Every tool is tested for actions above limits, without approval or on the wrong account.
- Jailbreak attempts include role-play, multi-step conversations and other languages.
- Tests try to push the agent into promises outside your policy.
- Hand-off summaries are tested for planted instructions and false identity claims.
- Rate limits, spend caps and cost alerts are tested with high-volume and looping inputs.
- Tests confirm that the agent checks identity before it shares account details.
- Fixes rely on permissions and approvals, not on prompt wording alone.
- Every fix is retested and added to an automated regression set.
- The next test date is set, along with triggers for model, prompt and tool changes.
For the controls to put in place before launch, use our secure AI agent deployment checklist.
Sources
- 1.Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1), NIST, 2024-07-26 (opens in a new tab)
- 2.AI Risk Management Framework, NIST, 2023-01-26 (opens in a new tab)
- 3.OWASP Top 10 for LLM Applications 2025, OWASP GenAI Security Project (opens in a new tab)
- 4.LLM10:2025 Unbounded Consumption, OWASP GenAI Security Project (opens in a new tab)
- 5.OWASP Top 10 for Agentic Applications for 2026, OWASP GenAI Security Project, 2025-12-09 (opens in a new tab)
- 6.GenAI Red Teaming Guide, OWASP GenAI Security Project, 2025-01-22 (opens in a new tab)
- 7.MITRE ATLAS, MITRE (opens in a new tab)
- 8.Joint guidance on deploying AI systems securely, CISA, 2024-04-15 (opens in a new tab)
- 9.Guidelines for secure AI system development: secure deployment, National Cyber Security Centre (UK), 2023-11-27 (opens in a new tab)
- 10.Prompt injection is not SQL injection (it may be worse), National Cyber Security Centre (UK), 2025-12-08 (opens in a new tab)
- 11.Regulation (EU) 2024/1689 (Artificial Intelligence Act), Articles 55 and 113, EUR-Lex, 2024-07-12 (opens in a new tab)
