Skip to content

AI security

AI red teaming: a practical guide and checklist for AI agents

AI red teaming means testing an AI system the way an attacker would: you try to make it leak data, ignore its rules or misuse its tools, so you can fix the weaknesses before someone else finds them. This guide explains how it differs from a penetration test and a model benchmark, what to test in an AI agent, how to run a test step by step and what the report should contain.

By the KDS security engineering teamPublished 11 min read

Key takeaways

  • AI red teaming tests an AI system the way an attacker would, to find how it can leak data, break rules or misuse tools.
  • It adds to a penetration test: it targets the instructions, content and tools the model works with, not only code and servers.
  • For AI agents, the most serious findings usually involve actions, such as refunds, data exports or messages sent without approval.
  • Retest after every change to the model, prompts, tools or permissions, not only before launch.
  • OWASP, MITRE ATLAS and NIST AI 600-1 give you shared test categories and a common language for findings.
On this page

What AI red teaming means

Red teaming is a long-standing security practice: a team plays the attacker to show how a system can fail. For AI, the idea is the same, but the target is wider. Besides code and servers, the tester attacks the instructions and content the model reads and the actions it can take. In the rest of this guide we call this work attack testing.

NIST describes AI red teaming as a structured exercise that probes an AI system for flaws and vulnerabilities, often in a controlled environment and together with the people who build it (NIST AI 600-1 (opens in a new tab)). The output for a business is practical: a list of ways your AI can be pushed into harm, each with proof, a risk rating and a fix.

For an AI agent that can issue refunds, update your CRM or send email, the most serious findings usually involve actions, not words.

How it differs from a penetration test and a benchmark

Three kinds of testing are often confused. They answer different questions, and a live AI agent usually needs all three.

AI red teaming compared with a penetration test and a model benchmark
TopicPenetration testModel benchmarkAI red teaming
Main questionCan an attacker break in through code, servers or accounts?How well does the model score on a fixed set of tasks?Can someone push this AI system into harm in its real setting?
What is testedApplications, APIs, networks and cloud setupThe model on its ownThe whole system: model, prompts, data sources, tools and hand-offs
Typical methodScanning and manual exploitation of known weakness typesStandard test sets and automatic scoringHostile conversations, poisoned content and tool misuse, both manual and automated
ResultFindings with severity and fixesA score or a pass rateFindings with evidence, severity, fixes and a retest
AI red teaming compared with a penetration test and a model benchmark

Benchmarks and the provider's own safety testing cover the model, not your system prompt, knowledge base, tools or permissions. Those parts are yours to test.

A penetration test is still needed for the application, APIs and cloud setup around the AI. Joint guidance on deploying AI systems securely (opens in a new tab), from cyber agencies in the United States, the United Kingdom, Australia, Canada and New Zealand, advises engaging external security experts to audit and penetration-test AI systems before deployment. For a side-by-side view, read our comparison of penetration testing and AI red teaming.

What to test in an AI agent

Start from what the agent can read and what it can do. The categories below follow the main risks in the OWASP Top 10 for LLM Applications (opens in a new tab) (LLM stands for large language model) and the OWASP list for AI agents.

Prompt injection

Prompt injection means tricking an AI with hidden instructions. In a direct attack, the user types them. In an indirect attack, they hide in content the agent reads: an email, a PDF, a web page, a support ticket or a record returned by a tool. It is LLM01, the first risk on the OWASP list. Test every input channel, not only the chat window. Our prompt injection guide covers the attack patterns in detail.

Data leakage

Check whether the agent reveals what it should keep private: its system prompt (the hidden instructions that set its role), other customers' records, internal notes, documents from restricted folders or keys stored in its configuration. OWASP lists sensitive information disclosure and system prompt leakage as separate risks. If the agent has memory, test whether data from one conversation appears in another.

Unsafe or unauthorized tool actions

Try to make the agent issue a refund above its limit, change an account's email address, delete a record, export data or message an outside address without approval. OWASP calls the root cause excessive agency: too many tools, too many permissions or too much autonomy. Its agent list adds tool misuse as well as identity and privilege abuse. Our guide to AI agent permissions explains how to set limits.

Jailbreaks and policy bypass

A jailbreak persuades the model to ignore its own rules. Testers use role-play, long multi-step conversations, other languages, encoded text and requests split into harmless-looking parts. The goal is to learn whether your business rules hold under pressure.

Hallucinated commitments

A model can state false facts with confidence, which OWASP groups under misinformation. For a customer-facing agent, the damage often comes from promises: a discount that does not exist, a refund outside your policy, a delivery date nobody agreed to or advice the agent is not allowed to give. Test whether a persistent customer can talk the agent into a commitment, and whether its answers stay tied to your approved sources.

Abuse of the hand-off to a person

A hand-off to your staff is a safety feature, but attackers can use it too. Test whether someone can plant instructions in the summary your staff read or make the summary present a false identity as fact. Also test whether a flood of hand-off requests can overload your team. OWASP lists human-agent trust exploitation among the risks for AI agents.

Denial of wallet and cost abuse

Every model call costs money. The OWASP risk unbounded consumption (opens in a new tab) includes denial of wallet: attackers send large volumes of requests to run up your AI usage costs. Test very long inputs, loops between tools, repeated retries and voice calls held open. Then check that rate limits, spend caps and alerts react as planned.

Privacy

Check what personal data the agent collects, repeats, stores in transcripts and sends to other systems. Test whether it confirms a caller's identity before it shares account details. The NIST Generative AI Profile suggests red teaming to check whether a system reveals personal, confidential or sensitive information.

How to run an AI red teaming project step by step

A useful test follows the same discipline as any security project: a written scope, a plan, careful testing and proof that the fixes work.

  1. Agree on scope and rules. Write down which agent, channels, tools and data are in scope and which environment you test. A test copy connected to test accounts is safest. Also agree what testers must not do (real payments, contact with real customers) and whom to call if something breaks.
  2. Build a threat model. List who might attack (customers, outsiders who send email, insiders), what they want (data, money, free services, damage to your reputation) and every path into the agent: user messages, files, retrieved documents, tool outputs and other agents.
  3. Write a test plan. Turn each threat into test cases with a clear pass or fail condition. Check coverage against the OWASP lists and MITRE ATLAS (opens in a new tab), a public knowledge base of attacker tactics, techniques and case studies for AI systems.
  4. Test by hand. Skilled testers chain steps, change approach when a defense holds and combine weaknesses, for example an injection that leads to a data export. Chained attacks are hard to find with scripts alone.
  5. Add automated testing. Tools send large sets of attack prompts in many variations and languages and record which ones succeed. The results become a regression set you can run again after every change. NIST notes that AI-led testing can be more cost-effective than human testers alone, and that each approach may find different kinds of harm.
  6. Rate each finding. Rate severity by impact and likelihood. Impact depends on what the agent can reach: an off-topic reply is minor, a tool call that moves money is not. Model outputs vary, so record how many attempts succeeded out of how many tries.
  7. Fix the design, not only the prompt. Stronger instructions help, but attackers can often talk around them. Prefer controls outside the model: narrower permissions, approval steps, code that checks tool inputs and separate data access for each user. The UK's National Cyber Security Centre (NCSC) recommends non-AI controls (opens in a new tab) that limit what the system can do.
  8. Retest and keep the tests. Confirm each fix with the original attack and its variants, then add the cases to your regression set.

When we run attack testing for clients, we follow these steps with a written scope and rules agreed in advance. We test in your environment on your own accounts and combine manual testing by security engineers with automated attack sets. Each finding comes with a reproducible example, a risk rating and a fix, and we retest after your team applies the fixes. See how our AI security testing works.

Frameworks to map your tests to

Frameworks give you shared names for risks and help you show auditors and customers what you covered. These are the most useful for AI agents:

  • OWASP Top 10 for LLM Applications: ten risk categories. The 2026 edition keeps prompt injection first and moves excessive agency up to third. Our guide to the OWASP Top 10 for LLM applications explains each one.
  • [OWASP Top 10 for Agentic Applications](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/): published in December 2025 for AI agents, with risks such as identity and privilege abuse, agent goal hijack and tool misuse. Our guide to the OWASP agentic top 10 explains each one.
  • [OWASP GenAI Red Teaming Guide](https://genai.owasp.org/resource/genai-red-teaming-guide/): published in January 2025, it covers four areas: model evaluation, implementation testing, infrastructure assessment and runtime behavior analysis.
  • MITRE ATLAS: tactics, techniques, mitigations and case studies of attacks on AI systems. It is useful for threat modeling and for describing how an attacker could use a finding.
  • NIST AI RMF and NIST AI 600-1: the voluntary AI Risk Management Framework (opens in a new tab) from January 2023 and its Generative AI Profile from July 2024. The profile suggests red teaming to test resilience against prompt injection and other attacks, and to check for leaks of personal or confidential data.
  • Guidelines for secure AI system development: published by the NCSC with CISA in November 2023, they advise releasing AI systems only after security evaluation such as benchmarking and red teaming (opens in a new tab).

What the EU AI Act says about adversarial testing

The EU AI Act names adversarial testing for one group: providers of general-purpose AI models with systemic risk. Article 55(1)(a) of the AI Act (opens in a new tab) requires them to conduct and document adversarial testing of the model to identify and reduce systemic risks. The rules for general-purpose AI models have applied since August 2, 2025.

A company that builds a support or sales agent on top of such a model is usually not the model provider. The provider's testing also does not cover your prompts, data or tools, so your own attack testing is still how you find risks in your setup. This is general information, not legal advice.

What a good AI red teaming report contains

A report should help your leadership decide what to fix first and let your developers reproduce every issue. Look for these parts:

  • A plain-language summary: the overall risk, the 3–5 most important findings and what to fix first.
  • Scope and method: the agent version, model, prompts, tools and data sources tested, the dates, the test environment and what was not tested.
  • Each finding in full: a clear title, severity, the affected component, the exact inputs and outputs, steps to reproduce, the success rate over repeated tries, the business impact and a recommended fix.
  • Framework mapping: the OWASP or MITRE ATLAS category for each finding, which helps with audits and customer security questionnaires.
  • Controls that held: the attacks that failed, so you know which defenses work.
  • Retest results: the status of each finding after your fixes.
  • The test cases: the attack prompts and scenarios, handed over so your team can run them again.

How often to test an AI agent

One test before launch is not enough. The model, the prompts and the tools change over time, and each change can undo a defense. Plan tests at these points:

  • Before launch, with enough time to fix the findings and retest.
  • After a model change, including a new version of the same model. The joint guidance advises a full evaluation, security tests included, before you redeploy an updated model.
  • After a change to prompts, tools or permissions, such as a new integration, a new data source or a new channel like voice.
  • On a schedule, for example every quarter for agents that handle payments or personal data, and at least once a year for the rest.
  • After an incident or a new public attack technique that could affect your design.

Between full tests, run your automated regression set after every release. It catches known attacks that start working again after a change.

AI red teaming checklist

Use this list to plan your first test or to review a testing proposal.

  • A signed scope names the test environment, test accounts, rules and emergency contacts.
  • A threat model lists the attackers, their goals and every input the agent reads.
  • Direct and indirect prompt injection are tested on every input channel.
  • Tests try to extract the system prompt, other users' data, personal data and stored keys.
  • Every tool is tested for actions above limits, without approval or on the wrong account.
  • Jailbreak attempts include role-play, multi-step conversations and other languages.
  • Tests try to push the agent into promises outside your policy.
  • Hand-off summaries are tested for planted instructions and false identity claims.
  • Rate limits, spend caps and cost alerts are tested with high-volume and looping inputs.
  • Tests confirm that the agent checks identity before it shares account details.
  • Fixes rely on permissions and approvals, not on prompt wording alone.
  • Every fix is retested and added to an automated regression set.
  • The next test date is set, along with triggers for model, prompt and tool changes.

For the controls to put in place before launch, use our secure AI agent deployment checklist.

Sources

  1. 1.Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1), NIST, 2024-07-26 (opens in a new tab)
  2. 2.AI Risk Management Framework, NIST, 2023-01-26 (opens in a new tab)
  3. 3.OWASP Top 10 for LLM Applications 2025, OWASP GenAI Security Project (opens in a new tab)
  4. 4.LLM10:2025 Unbounded Consumption, OWASP GenAI Security Project (opens in a new tab)
  5. 5.OWASP Top 10 for Agentic Applications for 2026, OWASP GenAI Security Project, 2025-12-09 (opens in a new tab)
  6. 6.GenAI Red Teaming Guide, OWASP GenAI Security Project, 2025-01-22 (opens in a new tab)
  7. 7.MITRE ATLAS, MITRE (opens in a new tab)
  8. 8.Joint guidance on deploying AI systems securely, CISA, 2024-04-15 (opens in a new tab)
  9. 9.Guidelines for secure AI system development: secure deployment, National Cyber Security Centre (UK), 2023-11-27 (opens in a new tab)
  10. 10.Prompt injection is not SQL injection (it may be worse), National Cyber Security Centre (UK), 2025-12-08 (opens in a new tab)
  11. 11.Regulation (EU) 2024/1689 (Artificial Intelligence Act), Articles 55 and 113, EUR-Lex, 2024-07-12 (opens in a new tab)

Get a free 30-minute assessment

Walk us through the AI tools and agents you use or plan to launch. We'll point out the biggest risks and the first fixes, in plain words.

Frequently asked questions

What is AI red teaming?

AI red teaming is testing an AI system the way an attacker would. Testers try to make a chatbot or AI agent leak data, ignore its rules, misuse its tools or make promises it should not make, then report what worked, how serious it is and how to fix it. For AI agents the focus is on actions, because a manipulated agent can change records, send messages or move money.

How is AI red teaming different from a penetration test?

A penetration test looks for weaknesses in applications, APIs, networks and cloud setup. AI red teaming targets what is new in an AI system: the instructions and content the model reads and the tools it can use. The two complement each other. A live AI agent usually needs both, because an attacker can use either path.

How long does an AI red teaming test take?

A focused test of one chatbot or AI agent takes 1–2 weeks in most cases, including the report. Agents with many tools, several channels or multi-agent workflows take longer. Plan time after the test for fixes and a retest, so the findings are closed before launch.

How often should you red team an AI agent?

Test before launch and after every change to the model, prompts, tools or permissions. Also test on a regular schedule, for example every quarter for agents that handle payments or personal data. Between full tests, rerun an automated set of known attacks after each release to catch problems that come back.

Can automated tools replace manual AI red teaming?

Not fully. Automated tools send large numbers of attack variations quickly and work well for regression testing. Skilled testers find chained attacks, adapt when a defense holds and judge the business impact. NIST notes that human and AI-led testing may find different kinds of harm, so a sound program combines both.

Does the EU AI Act require AI red teaming?

Article 55 of the AI Act requires adversarial testing from providers of general-purpose AI models with systemic risk, and these rules have applied since August 2, 2025. Most businesses that deploy AI agents are not in that group. Frameworks such as the NIST AI RMF and many customer security reviews still treat testing as good practice. This is general information, not legal advice.

Services and use cases

  • Customer support lead calmly reviewing a short list of escalated tickets at her desk

    Use case

    AI customer support

    AI customer support agents that resolve order, return and account questions from start to finish and hand complex cases to your team with full history.

    See how the AI support agent works
  • Service

    AI security and red teaming

    AI security consulting for AI agents and LLM apps: red teaming, prompt injection testing, shadow AI discovery and EU AI Act and ISO/IEC 42001 readiness.

    Explore AI security consulting
  • Service

    AI agent development

    Custom AI agent development for support, sales, front desk and back-office work. Voice and chat agents built and attack-tested by a cybersecurity team.

    Explore AI agent development

Free 30-minute assessment

Find the one workflow worth automating first.

Tell us how your team works. We'll come back with two or three AI opportunities, the risks to watch and a rough payback estimate. No obligation.

  • A senior engineer replies within one business day
  • We can sign an NDA before you share details
  • No fixed packages, every quote tailored to you
What can we help with?
About your company

Company size

When would you like to start?

How can we reach you?

Encrypted in transit · read only by our team · never sold

Free 30-minute AI assessmentGet it →

Before you go

Find out where AI can save your team time

Book a free 30-minute assessment. A senior engineer reviews one workflow with you and sends back the opportunities, the risks and a rough payback estimate.