What does securing a Claude agent in production mean?

Once an agent executes actions with real access, with no human reviewing each step, two questions take priority: how far is it allowed to reach, and what happens when someone tries to turn it against its operator? Accurate answers remain necessary, and they say nothing about either point.

Securing a Claude agent in production means limiting its permissions to what each task requires, keeping secrets out of its context, logging every action so it can be audited, and reserving high-impact decisions for a named human. This guide covers those workstreams, using Koneetiv's LOOP™ governance as the grid for classifying risk.

The gap between adoption and oversight is documented. According to Deloitte (State of AI in the Enterprise, 2026 edition), only 21% of organisations have a mature governance model for autonomous agents. For companies of every size, and even more so in regulated sectors, where an agent's action may have to be justified to an auditor, that gap surfaces at the moment of going live.

What an agent changes for security

A poorly tuned chat assistant produces a wrong answer, which a person reads before doing anything with it. An agent can update a record or call an API, and chain several actions of that kind inside a single task. The risk is then measured by the reach of its access.

What OWASP calls excessive agency

The OWASP Top 10 for LLM Applications, in its 2026 edition published in August 2026, moves Excessive Agency up from sixth to third place, behind prompt injection and sensitive information disclosure. The entry covers damaging actions triggered by unexpected, ambiguous or manipulated model output, and names the causes: excessive functionality, permissions or autonomy. Those excesses come from design decisions, taken when the agent is wired up, long before an attacker takes an interest.

An agent-specific reference since December 2025

In December 2025 the OWASP GenAI Security Project published a Top 10 for Agentic Applications. It covers agent goal hijack through injected instructions (ASI01), misuse of otherwise legitimate tools (ASI02), identity and privilege abuse (ASI03) and supply chain vulnerabilities, connectors included (ASI04). These public references give IT and the CISO a shared vocabulary for rating an agent's risks before it is connected to anything.

Least privilege applied to agent permissions

During a prototype, it is tempting to let the agent use the access of whoever configured it, or a broad service account. In production the rule flips: an agent holds only the access its task requires, checked action by action.

Scope access at the level of the action

An agent drafting customer replies needs to read the customer file and write a draft in the messaging tool. It needs neither write access to the billing database nor admin rights on the CRM. That breakdown forces you to document, for every possible action, what it reads and what it writes. The work feels tedious at the start, and it becomes the only reliable baseline the day the agent gains a new capability.

Enforce authorisation in the target system

OWASP's Excessive Agency entry recommends implementing authorisation in application logic, without letting the model decide whether an action is allowed. In practice, an instruction in the system prompt never replaces a permission denied at the API. When the agent acts on behalf of a user, its actions should run with that user's rights, and never with those of a technical account that can see everything.

Isolate the execution environment

An agent that runs code, executes commands or handles files should work in a sandbox with restricted network and file system access. The OWASP agentic Top 10 lists unexpected code execution among its risks (ASI05): if the agent is tricked, isolation determines what it can reach. A dedicated container per agent, with no access to other services' credentials, caps the blast radius of an incident.

Rotation and lifetime of credentials

A permanent access token stays usable for as long as its leak goes unnoticed. Regular credential rotation narrows the exploitation window if a secret leaks, for instance through a poorly filtered log or a code repository. For an agent running continuously, we recommend short-lived credentials, renewed automatically.

Secrets and credentials: what must never go through a prompt

A Claude agent receives its instructions and context as text. Pasting an API key or a password into a prompt, or into a document the agent will read, therefore hands that secret to anything that can influence the agent.

Exfiltration through prompt injection

An agent that reads external content, an email or a web page, may meet text designed to hijack its behaviour. OWASP keeps prompt injection first in its 2026 Top 10 (LLM01) and separates direct injection, carried by the user's message, from indirect injection, hidden in third-party content. Anthropic does not claim the problem is solved: in research published on 24 November 2025 on Claude's browser use, the company writes that no browser agent is immune to prompt injection. A secret sitting in an agent's context is exposed to every source that agent reads.

Keep secrets in a dedicated vault

In its 2026 edition, OWASP replaced its system prompt leakage entry with "Hidden Context Exposure" (LLM08), which covers all the hidden context placed in front of the model: system instructions, retrieved policy text, tool schemas, configuration directives. The guidance is blunt: no credentials, secrets or security-critical configuration in system prompts or hidden context, and that context should not be the primary mechanism for controlling model behaviour. The prompt injection entry points the same way, asking for credentials and state-changing capability to be held in application code, outside the model. Secrets therefore live in a vault, the agent gets temporary access when it needs it through an orchestration layer that handles authentication, and it never sees the raw value. That separation also protects against an agent copying a credential into its own output, by mistake or under manipulation.

Logging: reconstructing what the agent actually did

Monitoring an agent's answers is not enough to audit it. After an incident, you need to establish exactly what it did, with which access and on which data.

What an agent log should contain

A usable audit trail keeps the requested task, the identity and access level used, the sequence of tool calls with their parameters, and the observed result. Each entry is timestamped and stored out of the agent's reach, so the agent can never rewrite its own record. That log is also what lets you reclassify an action, based on behaviour observed over time.

What the AI Act requires, and from when

For high-risk systems, Article 12 of the AI Act requires them to technically allow automatic recording of events over their lifetime, and Article 14 requires that they can be effectively overseen by people. The Digital Omnibus on AI, Regulation (EU) 2026/1744, in force since 27 July 2026, moved those obligations to 2 December 2027 for Annex III systems, which include recruitment, education and creditworthiness assessment. An agent that screens applications or prepares a credit decision falls within that scope. A log designed now avoids a rushed rebuild later, and it already serves as evidence in an ISO 42001 audit. Our page on AI Act compliance for companies sets out the full timeline.

The 4 LOOP™ trust zones applied to security risk

LOOP™ governance (Living Oversight & Operations Protocol) assigns every possible action of an agent to a trust zone, defined before go-live. The zone is attached to each action: the same agent can generate a report in the green zone and have another of its actions blocked in the black zone.

Green zone: autonomous execution

The agent acts without prior approval and the result is logged. This zone holds reversible, low-risk, high-volume tasks, such as first-level ticket triage or standard report generation. From a security angle, this zone implies narrow access: a mistake there costs little and is quickly corrected.

Orange zone: human approval before execution

The agent prepares a complete, reasoned recommendation, and a designated person approves it before the action goes out. LOOP™ places supplier payments above a threshold, sensitive customer replies and contract changes in this zone. It is the natural home for write operations that commit the company.

Red zone: mandatory escalation

The agent stops, documents what it understood and what it is missing, then alerts the owner named in the register, who must decide within 4 hours. No automatic action is possible. LOOP™ places unforeseen ambiguous situations and unexpected sensitive data in this zone.

Black zone: immediate block and CISO alert

The black zone covers anything outside the agent's remit: access to data beyond its authorised classification, requests to bypass guardrails, prompt injection attempts. Execution stops, the CISO is alerted and the full audit trail is switched on. This is where governance and security overlap most clearly, since LOOP™ places several forms of attack in this zone.

Escalation and the living register

LOOP™ defines three escalation levels: information, which notifies without blocking; validation, which requires human approval before the action; and blocking, which halts everything, alerts the CISO and activates the audit trail. Every agent is recorded in a living register that documents, among other things, its owners (business, technical, escalation contact, CISO approver), the data it accesses and the scope of its authorised actions. We recommend reviewing an action's classification whenever the agent gains a connector or a write capability. Our article on trust zones for classifying AI agents details the method, and the LOOP™ governance page presents the protocol.

Where to start securing your agents

We advise starting with an inventory of the agents already running, informal ones included, then classifying their actions by their real access, as it appears in the target systems. S&P Global reported in March 2025 that the average organisation scrapped 46% of its AI proofs of concept before production, with respondents citing cost, data privacy and security risks among the top obstacles. Raising security questions at scoping avoids discovering them at the go-live committee.

Test, then audit the connectors

Two checks complete permissions, secrets and logging. The first is to actively test the agent's resistance to hijacking attempts, which our article on red teaming Claude AI agents covers. The second concerns the MCP servers through which the agent reaches your systems: each one adds a surface to audit, detailed in our article on auditing MCP servers in production. For agents built with Claude Code, our guide to connecting Claude Code to enterprise systems with MCP covers permission and catalogue settings on the workstation side.

Measure where you stand

Our AI maturity assessment places your company on 6 axes, including governance and industrialisation, in three minutes, using the Koneetiv framework (2026 edition). For organisations that want continuous oversight, Claude Cockpit records every agent in the LOOP™ living register, with its trust zones, escalation thresholds and supervision dashboards.