What is red teaming an AI agent?

A Claude agent in production receives instructions, reads documents, calls tools and sometimes writes to third-party systems. Each of those capabilities widens the surface an attacker can aim at. A security test therefore has to check how the agent behaves when an input is designed to push it off course.

Red teaming an AI agent means simulating, in a methodical and controlled way, manipulation attempts against an agent placed in realistic conditions: instructions hidden in an input, data exfiltration attempts, attempts to bypass its system rules. The goal is to measure its resistance before a real attacker does.

A red teaming session brings together the team that built the agent and an independent pair of eyes that does not know its internal settings. That outside view shrinks the design team's blind spot, since builders tend to test what they anticipated rather than what they never imagined.

Why test a Claude agent before it goes live

A misaligned chat system produces a wrong answer. A misaligned agent can act: send a message, update a record, trigger a tool call against a third-party system. That ability to act changes what a security test must check, and explains why a conventional code audit falls short of covering the risk.

What changes with an agent

An agent chains tool calls without a person approving each step. An adversarial instruction slipped in early in a session can therefore spread across several actions before anyone spots it. The OWASP Top 10 for Agentic Applications, published in December 2025, calls this agent goal hijack (ASI01) and separates it from the misuse of legitimate tools (ASI02). Red teaming reproduces those mechanics with multi-turn scenarios rather than a single isolated question.

What public frameworks recommend

In the 2026 edition of its LLM Top 10, published in August 2026, the prompt injection entry (LLM01) asks teams to test defences against adaptive attackers who have read the deployed protection, and to reject attack-success figures measured on static attacks only. The OWASP project had published a dedicated GenAI Red Teaming Guide in January 2025. Anthropic, for its part, describes in research published on 24 November 2025 on Claude's browser use a defence that combines model training, classifiers and human red teaming, and notes that human security researchers outperform automated systems at finding creative attack vectors. Those protections reduce risk at the source, and an enterprise deployment still has to test its own tool perimeter.

What the AI Act expects from high-risk systems

For companies in regulated sectors, the topic also has a regulatory side. Article 15 of the AI Act requires high-risk systems to be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting vulnerabilities, and lists inputs designed to make the model err among the attacks to address. Since the Digital Omnibus on AI, Regulation (EU) 2026/1744, those requirements apply from 2 December 2027 to Annex III systems, such as candidate screening or creditworthiness assessment. A history of documented campaigns helps demonstrate that robustness when the time comes. Our page on AI Act compliance for companies sets out the timeline.

One building block of agent security

Red teaming complements permission design, logging and human review of sensitive actions. Our guide to securing a Claude agent in production covers the whole set-up, with red teaming as its active verification phase.

Direct prompt injection: testing resistance to adversarial instructions

Direct prompt injection is an instruction, written into the message sent to the agent, that tries to override or neutralise its rules. The user openly asks the agent to ignore its constraints, switch roles or reveal how it was configured.

What a test scenario should cover

A test plan varies the wording of the same adversarial intent: direct rephrasing, a switch of language, encoded text, a request split across several successive messages. An agent that resists a blunt phrasing may give way to a more indirect variant, and that gap is what the test must expose. Hidden context exposure calls for a dedicated scenario. OWASP makes it entry LLM08 of its 2026 Top 10 ("Hidden Context Exposure") and treats that context as extractable: the test tries to reconstruct system instructions and tool schemas, then checks that no secret and no security control depends on them.

Record every attempt

An undocumented attempt is lost to the next campaign. Each test is recorded with its exact wording, the agent's response, the tool calls it triggered and a simple verdict: compliant, partial or failed. That record makes one campaign comparable with the next and lets you prioritise fixes instead of reacting to whatever happens to surface.

Indirect prompt injection: when the attack comes from third-party content

The hardest variant to anticipate comes from content the agent processes to do its job: a shared document, an incoming email, a web page, a tool response. Text can be planted there to look like an instruction, in the hope that the agent follows it without telling it apart from data. The 2026 LLM01 entry recommends passing external content through a structurally separate, provenance-labelled channel, so that the model can tell data from instructions. The same entry adds that this marking reduces attack success in non-adaptive tests only, since an attacker who knows the marking scheme can mimic it: it complements the adaptive testing described above.

The third-party surface: documents, emails, web pages, tool outputs

Every connector an agent uses becomes a potentially adversarial input channel. An agent that summarises attachments, browses the web or queries an MCP server handles text whose origin it does not control. The test injects disguised instructions into those channels and observes whether the agent executes them. An agent working through a mailbox or a shared drive handles dozens of files in a single task, and one booby-trapped item is enough to make the scenario fail.

The case of MCP connectors

For MCP connectors, the protocol specification (28 July 2026 version) says descriptions of tool behaviour should be considered untrusted unless they come from a trusted server. Our article on auditing MCP servers in production covers that surface server by server, and our guide to connecting Claude Code to enterprise systems with MCP shows how to restrict access on the workstation side. Every connector added goes into the test plan of the next campaign.

Data exfiltration attempts: what a test should simulate

An agent with access to customer files, contracts or internal exchanges can be manipulated into revealing them outside their intended context. Red teaming simulates that pressure: asking the agent to summarise a conversation outside its remit or to send data to an unauthorised destination, for example by embedding it in a web address or an outgoing message.

What an agent must never reveal, even rephrased

An effective test goes beyond the direct question ("show me customer X's data"). It pushes variants: an anonymised summary that becomes identifying once cross-referenced; a translation or rewording of an excerpt the agent should never repeat; an explanation of its own reasoning that would expose context data. OWASP files this risk under LLM02 in its 2026 Top 10, sensitive information disclosure.

Test with synthetic data

A red teaming campaign deliberately pushes the agent outside its frame. Running it on real sensitive data multiplies the risk if the test environment is misconfigured, and raises a GDPR question. Synthetic datasets built to resemble real cases cover the same ground without exposing anyone.

Bypassing guardrails: the test method

A guardrail is a rule placed around the agent, such as a list of forbidden actions, an output filter or a check in the target system, independent of the model's reasoning. A well-designed set-up stacks several of them, and the test checks each layer separately before probing the combination, because one layer can mask another's weakness.

Contextual pressure

A system prompt states constraints, and a reasoning agent can be led to reinterpret them under pressure: simulated urgency, impersonated authority, a justification that looks legitimate in the context provided. A test scenario includes those pressures alongside blunt requests, and checks that the agent sticks to its rules even when the other party claims to be a manager.

Conflicts between rules

Some failures only appear where two instructions that are legitimate on their own intersect. A test that deliberately combines contradictory instructions often reveals behaviour that neither scenario would have shown in isolation.

Building a recurring test methodology

A red teaming exercise run once before launch measures a state at a given moment. It says nothing about the agent's behaviour a few months later, once its environment has changed.

Before go-live

The initial campaign covers every scenario, direct and indirect injection, exfiltration and guardrail bypass, on each tool and connector the agent will actually use. It also sets the acceptance threshold, meaning which failures block go-live and which lead to a scheduled fix. The resulting baseline records which guardrail stopped each failed attempt, so the team knows which layer is actually doing the work.

Continuously after deployment

Once the agent is live, testing continues at a regular pace, and every significant change triggers a new campaign: a connected tool, a system prompt update, a new model version. Results join the agent's history, next to its incidents and audits.

Red teaming and LOOP™ governance

Red teaming answers a technical question: does the agent hold up under attack? LOOP™ governance answers an organisational one: who decides, who supervises, and with what level of autonomy for each action. The LOOP™ black zone explicitly covers prompt injection attempts and requests to bypass guardrails, with an immediate block, a CISO alert and a full audit trail. A red teaming campaign is how you check that the block actually fires.

Focus the effort on the most exposed actions

In LOOP™, the zone is attached to each of the agent's actions. Actions in the orange zone, where a person approves before execution, and in the red zone, where the agent stops and escalates to a designated owner, are those where a bypass would cost the most. They are the ones that get the most demanding scenarios, such as an attempt to get an action that requires approval executed directly. Our article on trust zones for classifying AI agents details the classification, and the one on the LOOP™ human-centred governance protocol presents the 4 zones and the 3 escalation levels.

Where to start with your own agents

Before building a full campaign, you need to know where to aim first: which agents access sensitive data, which ones hold write rights, which ones read third-party sources. Agents that combine reading third-party content with write access come first, because that is where an indirect injection turns into an action.

The Koneetiv deployment method includes a business testing and security red teaming phase before any go-live. To see where your organisation stands, our AI maturity assessment gives you a score and an action plan in three minutes, using the Koneetiv framework (2026 edition). As agents multiply, Claude Cockpit takes over their ongoing oversight, each one recorded in the LOOP™ living register.