Agent Security Architecture

The Action Boundary Gate

AI agents that authorize their own actions are the most exploitable systems in enterprise AI. The fix is architectural, not behavioral. A deterministic permission layer between model proposal and real-world execution changes the security posture entirely.

Arjun Jaggi  ·  September 19, 2026  ·  14 min read
0
production AI agents that should be authorized to approve their own actions
3
distinct failure modes that converge at the action boundary: prompt injection, hallucination, bad reasoning
1
architectural separation that defines a secure agentic pipeline from an exploitable one

The Problem Sitting in Your Agent Pipeline

Your CISO owns this, even if they do not know it yet. Every AI agent you have deployed, piloted, or are evaluating has the same structural property: the model that decides what to do is also the system that does it. The reasoning and the execution are coupled. There is no checkpoint between the model's conclusion and the action it triggers in your environment.

This matters because the model's confidence is not authorization. A model that reasons its way to "I should write this file" or "I should call this API endpoint" has done nothing more than produce a proposal. Whether that proposal is permitted, in this context, for this principal, against this resource, is a question that the model cannot reliably answer about itself. The model is subject to prompt injection, hallucination, and reasoning failures that can produce confident proposals for actions that should never execute. Greshake et al. documented the shape of this attack surface in 2023, showing how indirect prompt injection turns trusted agents into vectors for attacker-controlled actions in production systems. [1] The attack does not require the model to be defective. It requires only that the model trust its input context more than it should.

The correct architecture inserts a deterministic control between what the model proposes and what the system executes. That control is the Action Boundary Gate. It does not ask the model to double-check its reasoning. It evaluates the proposed action against a permission policy using logic that is not subject to the same failure modes as the model itself.

Coined Term

Action Boundary Gate (ABG): the deterministic control layer that sits between a model's proposed action and its real-world execution, evaluating the proposed action against a permission policy without re-invoking the model. The ABG is the enforcement point for the principle that model confidence is not authorization. It operates on action metadata (type, target, scope, principal) rather than on model reasoning, and its decisions are recorded in an immutable audit log independent of the model's context window.

The concept is not new in security architecture. Pre-execution authorization at a trust boundary is a standard pattern in privileged access management, in network policy enforcement, and in operating system syscall interception. What is new is the context in which it is being systematically skipped: enterprise AI agent pipelines, where the pressure to demonstrate speed has consistently outpaced the pressure to demonstrate correctness.

The Architecture

Action Boundary Gate Architecture
MODEL proposes action proposal ACTION BOUNDARY GATE action type check principal validation scope enforcement PERMISSION POLICY declarative, versioned ALLOW EXECUTION real-world effect BLOCK + AUDIT LOG immutable record AUDIT LOG all allowed actions gate trust boundary TARGET SYSTEM file, API, shell, tool BLOCK

The architecture has two invariant properties. First, the ABG is evaluated before execution, never after. A post-execution audit log is necessary but insufficient: it records what happened, it does not prevent it. Second, the ABG evaluates action metadata, not action reasoning. It does not ask whether the model's decision to act was sensible. It asks whether the proposed action is within the permitted scope for this agent, this principal, this resource, at this time. These are answerable by deterministic policy, not by model reasoning.

Coined Term

Execution Trust Split: the architectural separation between the system that proposes actions (the model, probabilistic, context-dependent, subject to injection and hallucination) and the system that authorizes actions (the gate, deterministic, policy-bound, not subject to the same failure modes). The Execution Trust Split is the structural property that distinguishes a secure agent pipeline from an exploitable one. In its absence, a single successful prompt injection can propagate from input to execution without any checkpoint.

Why the Model Cannot Be Its Own Gatekeeper

This is the failure mode that practitioners consistently underestimate. When teams ask "what could go wrong with our agent," they list the scenarios they can imagine. Prompt injection is on the list. Hallucination is on the list. Bad reasoning on edge cases is on the list. What is frequently absent is the recognition that all three of these failure modes share a common consequence structure: they produce a model that is confidently proposing an action it should not take, and without an ABG, that confidence is sufficient for execution.

Adding a second model call to "verify" the first model's proposal does not solve this. A second model call shares the same failure modes as the first. Indirect prompt injection that corrupts the context of the primary model is equally capable of corrupting the context of a verification model drawing from the same inputs. The Greshake et al. analysis of real LLM-integrated applications documents this class of cascading failure explicitly: the attack surface is the model's relationship to its input context, and adding more model calls to the same compromised context deepens the attack surface rather than narrowing it. [1]

A deterministic gate operating on action metadata breaks the chain. The injected payload that convinced the model to exfiltrate a file cannot convince the ABG, because the ABG does not read the payload. It reads the proposed action type, the target resource, and the requesting principal, and compares them against a static policy. That policy was authored by a human, versioned in source control, and is not reachable by the model's input context.

"Permission checks for AI agents have to happen at the action boundary, before execution. A model can propose a shell command, file write, network call, or tool invocation, but a separate deterministic control should decide whether that exact action is allowed. That separation is what keeps prompt injection or bad reasoning from automatically becoming a real-world side effect."
Michael Kantor, [email protected], HOL Guard. hol.org/guard/guides/ai-agent-security-layers [2]

Kantor's framing identifies the precise architectural principle at stake. The separation between model proposal and gate decision is not a performance tax or an operational overhead. It is the mechanism by which an entire class of failure modes is structurally prevented. Without that separation, every agent in your environment is as trustworthy as its worst possible input context. With it, the worst possible input context can still only produce a blocked proposal and an audit log entry.

Gate Decision Framework

Not all proposed actions require the same gate response. The appropriate gate behavior depends on two independent variables: the reversibility of the action and its blast radius within the enterprise environment. These two variables produce four gate-response categories, each with a different enforcement posture.

Action Category Reversibility Blast Radius Gate Response Audit Requirement
Read-only data access N/A Scoped to data asset Allow with log Access log entry
Scoped file write Reversible Workspace-contained Allow with log Write record with diff
Internal API call Varies by endpoint Service-contained Allow if in allowlist Request/response log
Shell command execution Varies by command Host-contained Allowlist-only, block all others Command + stdout/stderr
Cross-system write or mutation Partially reversible Enterprise-wide Notify owner, await veto window Intent + approver + outcome
External network call Not reversible Exfiltration risk Block unless explicitly allowlisted Blocked attempt log
Credential access or rotation Not reversible Enterprise-wide Block, alert security team Security incident record

This taxonomy composes with the Action-Consequence Classification framework introduced in Agent Autonomy Calibration, which provides a formal treatment of how reversibility and blast radius jointly determine the appropriate autonomy posture for any given action. The ABG is the enforcement mechanism for whatever autonomy level the ACC framework assigns. The frameworks are complementary: ACC classifies, ABG enforces.

Gate Decision Simulator

Action Boundary Gate Simulator

Failure Modes Without the Gate

Agent Security Incident Source Analysis
Indicative incident exposure by failure mode: ungated vs. gated pipeline Without ABG With ABG 0 25 50 75 100 85 8 Prompt Injection 70 65 Hallucination (wrong target) 65 58 Bad Reasoning (edge case) 80 15 Permission Drift 60 12 Cascading Agent Auth
Directional illustration. Distribution is based on practitioner observation across documented agentic AI incident classes. Categories are not mutually exclusive. Not derived from systematic survey data.

Three Enterprise Scenarios

Financial Services  ·  CISO

Incident Response Agent with Unrestricted Firewall Access

A financial services firm deployed an incident response agent with authority to modify firewall rules based on anomaly detection. The agent was not operating behind an ABG. A phishing email containing a prompt injection payload reached the agent's processing context through a connected email security tool. The injected payload instructed the agent to block an IP range matching the firm's own clearing infrastructure. The agent complied, triggering a four-hour outage affecting institutional clearing operations. The action was technically reversible. The downstream consequences were not. An ABG with a policy rule blocking IP ranges in the firm's protected CIDR space from agent-initiated modification would have blocked the proposed action and flagged it for the security team before execution.

Healthcare  ·  CTO

Data Extraction Agent with Overly Broad Read Scope

A healthcare technology team deployed a research assistant agent with read access to the clinical data lake for the purpose of answering questions about trial data. The agent's permission model was defined at deployment and was not enforced at the action boundary during runtime. A user query with an embedded indirect injection caused the agent to issue a bulk export query across patient records beyond the scope of the intended access. The ABG pattern would have evaluated the proposed query against the agent's declared data scope at query time, blocking any access beyond the registered dataset boundaries regardless of what the model concluded from its context.

Enterprise SaaS  ·  CIO

Code Deployment Agent Without Pre-Execution Authorization

A software development organization piloted an agentic CI/CD assistant with authority to merge pull requests and trigger deployments in response to review completion signals. The agent had no gate between its decision to deploy and the deployment action. A hallucination in the model's review logic caused it to interpret an incomplete code review as a completed one and initiated a deployment to the production environment. The ABG pattern, with a policy requiring a verified human approval signal before any production deployment action, would have blocked the proposed deployment and returned a notification requiring explicit authorization. The pilot was suspended after the incident. A gate-first architecture would have kept it running.

Risk Register

Risk 1. Prompt Injection Escalation

Indirect prompt injection via connected data sources (email, documents, external APIs) injects instructions into the agent's context that produce authorized-looking action proposals. Without an ABG, the injection payload executes. With an ABG, the proposed action is evaluated against policy regardless of the instruction source. Early signal: agent producing action proposals that reference resources outside its declared scope.

Risk 2. Permission Drift

Agents accumulate permissions over their operational lifetime through ad hoc exceptions granted during development and integration testing. Without an ABG enforcing a versioned policy at runtime, actual permissions diverge from intended permissions without detection. The drift is invisible until an incident surfaces it. Detection mechanism: periodic diff between declared ABG policy and observed agent action logs.

Risk 3. Cascading Agent Authorization

In multi-agent architectures, an agent that can spawn or direct sub-agents can propagate permissions to those sub-agents in excess of what each sub-agent was individually authorized for. Without an ABG at each agent's action boundary, the orchestrator's authorization becomes the effective authorization for the entire pipeline. The Multi-Agent Trust Propagation post covers this failure mode in detail. The ABG must be implemented per agent, not per pipeline.

Risk 4. Audit Gap

Without an ABG, the only audit record available is application-level logging, which is controlled by the same system that executed the action and may not record blocked intentions or failed authorization attempts. An ABG produces an audit record that is independent of the model and captures proposals as well as executions, giving incident responders the full picture of agent intent rather than only the subset of intent that reached execution.

Risk 5. Regulatory Non-Compliance

AI governance frameworks emerging across jurisdictions are converging on a requirement for human oversight of high-stakes AI actions, with traceability requirements for any AI-initiated action affecting regulated data or processes. An agent architecture without an ABG cannot produce the traceability record required to demonstrate that each consequential action was within authorized scope. The EU AI Act's requirements for high-risk AI systems impose documentation and traceability obligations that an ABG-based architecture satisfies structurally, while an ungated architecture satisfies only partially and retroactively.

Build vs. Buy vs. Configure

Build

Policy Engine and Audit Log

The declarative permission policy for your agents and the immutable audit log infrastructure are domain-specific and should be built internally. Off-the-shelf policy engines require mapping to your agent action taxonomy. Audit logs must be tamper-resistant and independently stored from agent runtime data.

Buy or Adopt

Gate Infrastructure and Enforcement Libraries

Open policy agent (OPA), service mesh policy enforcement, and tool-level interception libraries are mature. Evaluate existing agent security tooling before building interception logic from scratch. The HOL Guard approach [2] provides a reference architecture for gate placement in common agent frameworks.

Configure

Agent Framework Hooks

Most production agent frameworks expose pre-execution hooks or middleware interfaces that can host gate logic without modifying the core model pipeline. Configure these hooks before adding gate logic. Retrofitting gates into frameworks that do not expose clean hooks is significantly more expensive than configuring hooks at initial deployment.

Implementation Roadmap

Phase 1. Weeks 1 to 4

Action Inventory and Policy Draft

Enumerate every action category each deployed agent can take. Draft a declarative permission policy for each agent covering action type, target resource, and principal scope. Do not deploy the gate yet. Validate the policy against observed agent behavior using shadow mode logging. Go/no-go gate: policy covers 100% of observed action types with no gaps.

Phase 2. Weeks 5 to 10

Gate Deployment in Audit Mode

Deploy the ABG in audit-only mode: evaluate all proposed actions against policy and log the decision, but do not block. Observe the gap between policy intent and agent behavior. Tune the policy to eliminate false positives before switching to enforcement mode. Go/no-go gate: false positive rate below the team's defined threshold for at least two continuous weeks.

Phase 3. Weeks 11 to 16+

Enforcement Mode and Governance Integration

Switch the ABG to enforcement mode. Integrate the audit log with your security information and event management system. Establish a policy review cadence tied to agent capability updates. For agents operating under regulated frameworks, begin producing the traceability documentation the ABG enables. Success criteria: zero ungated agent actions in production environments.

Minimum Viable Team

A pilot ABG deployment for a single agent category requires a focused team. At the pilot stage: one Senior Platform Engineer who owns gate infrastructure and policy engine integration, one Security Architect (part-time) who owns policy authoring and threat modeling, one ML Engineer who owns the agent framework hooks and shadow mode instrumentation, and one product or program owner with AI literacy who owns the action taxonomy and stakeholder alignment. Scale-up adds a dedicated policy engineer and a security analyst for ongoing audit log review. The governance integration in Phase 3 typically requires legal or compliance involvement for regulated environments.

The Containment Thesis

The Contained Agent framework establishes the formal property that secure agentic architectures require at minimum one point of deterministic enforcement between model decision and real-world effect. The ABG is that enforcement point. Its absence is not a gap in tooling that will be filled by better models. It is a structural property of how the system is built, and it cannot be improved away by model upgrades. The agent that is one prompt injection away from executing arbitrary file writes today will still be one prompt injection away from executing arbitrary file writes with a more capable model tomorrow, unless the architecture changes.

The Action Boundary Gate is the architecture change. The Execution Trust Split it implements is the property that makes agentic systems trustworthy at enterprise scale. This is the missing layer that the AI agent security discussion consistently circles without naming. The Shadow AI Surface analysis in Shadow AI Surface identifies the inventory problem that precedes the gate problem: you cannot build gates for agents you do not know are running. Solve the inventory problem first. Then build the gates.

Executive Sign-Off Checklist
Does every agent in production have a documented action taxonomy?
Good: taxonomy exists, is versioned, and matches observed behavior. Red flag: "the agent can do what it needs to do" without an enumerated list.
Is there a deterministic gate between model proposal and action execution?
Good: gate is in enforcement mode with audit logging for all agents. Red flag: gate is "planned" or exists only for a subset of agents.
Is the gate policy stored independently of the agent runtime?
Good: policy is in version control, reviewed on a defined cadence, and not modifiable by the agent. Red flag: policy is inline in the agent code or modifiable at runtime.
Does the audit log capture blocked proposals as well as allowed executions?
Good: audit log is independent of agent runtime and captures intent before gate decision. Red flag: logging only captures successful actions.
Have you tested the gate against a simulated prompt injection in staging?
Good: injection test is part of the agent deployment checklist and results are documented. Red flag: "we trust the model to handle that."
Is the gate architecture documented for regulatory review?
Good: gate policy, audit trail, and override procedures are documented and accessible to compliance. Red flag: architecture exists but is not documented for audit purposes.
Do sub-agents in multi-agent pipelines have their own gates?
Good: each agent in the pipeline is gated independently, not relying on the orchestrator's authorization. Red flag: only the orchestrating agent has a gate.
Is there a defined process for policy updates when agent capabilities expand?
Good: policy update is a required step in the agent capability release process. Red flag: policy updates are ad hoc and not tied to the deployment process.

Excited about AI, innovation, and growth?

Start a conversation

References