This paper introduces the Agentic Attack Surface Framework (AASF), an original taxonomy and control architecture for securing enterprise AI agents. It is written for the leaders who must sign off on agent programs, the CISO who must defend them, the CTO who must build them, and the board that must price the risk of both.
The Agentic Attack Surface Framework, its ten exhibits, the four-layer control architecture, and the posture model are original contributions of this work, published under the license on the back cover. The constructs Action Blast Radius, Memory Drift, and Recall Boundary are drawn from the author's prior research and applied here to agent security, formally defined in the concept papers cited in the notes [7][8]. Every quantitative claim is cited to a named source or labeled directional, without exception.
Every generation of enterprise technology arrives with a security model borrowed from the generation before it, and every generation discovers, at some cost, where the borrowed model breaks. Client-server borrowed from the mainframe. Cloud borrowed from the data center. AI agents are now borrowing from all of them, and the fit is worse than it has ever been.
The reason is simple to state and hard to govern. An agent is the first enterprise system whose inputs are also its instructions. The document it retrieves, the email it reads, the tool output it parses, each one can carry a command, and the agent cannot reliably tell the difference. Researchers demonstrated this against real deployed applications within months of the first agent frameworks shipping [3]. Security teams are being asked to govern a system that takes orders from its own reading material.
This paper is my answer to the question I am asked most often by executive teams. Not whether to deploy agents, that decision has already been made by the market, but how to deploy them so that a single poisoned document cannot become a seven-figure incident. The framework in these pages is original work. The numbers in these pages are cited or labeled directional, without exception. I hold this work to the standard of the rooms it will be read in.
An AI agent is not a smarter chatbot. It is an autonomous system with persistent memory, live tool access, and the authority to act on behalf of the enterprise. That combination creates an attack surface that perimeter-era security was never designed to see. Five findings follow.
Agents collapse the boundary between data and instruction. Every document an agent retrieves and every tool output it reads is a potential command channel. Indirect prompt injection has been demonstrated against real deployed applications [3], not laboratory models. This is a present condition, not a forecast.
Exposure scales with tool access, not data footprint. One compromised agent with email, file, and SaaS reach can move data across all three in a single task sequence. Governance must bound the Action Blast Radius, the set of irreversible actions reachable before a human gate, before the program scales.
Six attack classes cover the agentic threat space. The AASF taxonomy maps each class to the system layer it enters through and the control layer that intercepts it. A gap in any one layer is reachable from a compromise earlier in the chain, which is why partial stacks fail.
Detection probability collapses after goal substitution. Once an agent's objective has been silently redirected, every later action looks legitimate in isolation. Controls inserted before tool invocation return the most protection per engineering hour, a result quantified in Exhibit 6.
Governed autonomy is reachable in one quarter. A four-layer control stack deployed across three gated phases takes a single-agent program from ad-hoc exposure to governed autonomy in roughly fifteen weeks, directionally, with a team of four to five practitioners.
Enterprise adoption of AI agents has outrun enterprise governance of them, and the gap is now measurable. Reported AI incidents reached 233 in 2024, a 56.4 percent increase over the prior year [2], while the average cost of a data breach reached $4.88 million [1]. Agents raise the stakes on both, multiplying how an incident can begin and what it can reach. The exposure is measured, not hypothetical. In one published benchmark a tool-using frontier model obeyed instructions hidden in retrieved content in 24 percent of 1,054 cases [6].
Traditional controls operate on syntax. Agent threats operate on semantics. A prompt injection attack arrives as a perfectly valid request whose content redirects agent behavior, and no firewall rule, rate limit, or schema check will catch it. The earliest public demonstrations planted instructions in web pages and documents that deployed assistants then obeyed [3][5].
Action Blast Radius. The total set of irreversible external actions an agent can execute before reaching a human approval gate, formalized as a scalar risk metric in the cited concept paper [7]. Every control in this paper exists to shrink it.
The structural change is shown in Exhibit 2. An agent sits inside one trust boundary with five distinct system layers, and an adversary can enter through any of them. The damage potential of a compromise is proportional to the agent's tool access, not the data it holds.
The Agentic Attack Surface Framework classifies threats by the system layer they enter through and the type of agency they exploit. Every agentic incident observed in practice resolves to one of these six classes, or to a chain of them.
Malicious instructions embedded in inputs, retrieved documents, or tool outputs redirect agent behavior without the orchestrator's knowledge. Demonstrated against deployed applications [3].
Manipulation of multi-step planning so the agent pursues an attacker objective. Hard to detect because each step looks legitimate in isolation.
Adversarial content written into persistent or working memory shapes all future reasoning, within and across sessions. The slowest attack and the hardest to unwind.
A compromised tool definition, output schema, or upstream API causes the agent to take unintended actions while believing it executed correctly.
The agent's legitimate tool access routes sensitive data to attacker endpoints. The agent acts as a compliant carrier, so no system compromise is required.
Agent-to-agent communication or tool chaining acquires permissions beyond those granted at session start. The signature risk of multi-agent architectures.
Recall Boundary. The set of information an agent may retrieve and act on within one task session, formally defined in the cited concept paper [8]. Exceeding it without authorization is the earliest reliable Class III signal.
Memory Drift. Gradual divergence between an agent's stored behavioral priors and its operational context, formally defined in the cited concept paper [8]. The gap drift opens is where Class III attacks live.
Each layer intercepts a different stage of the Exhibit 3 chain, and a gap in any layer is reachable from a compromise earlier in the chain. Traffic that clears Layer 1 is still watched by Layer 2, still scoped by Layer 3, and still bounded by Layer 4.
| Attack class | Primary layer | Earliest detection signal | Mitigation | Residual risk |
|---|---|---|---|---|
| I. Prompt Injection | Layer 1 | Anomalous instruction pattern in input | Classifier plus schema enforcement | Medium |
| II. Goal Hijacking | Layers 1 + 2 | Divergence from declared objective | Goal drift monitor plus human gate | Medium |
| III. Memory Poisoning | Layers 2 + 4 | Recall Boundary exceeded | Memory scoping plus access audit | High |
| IV. Tool Poisoning | Layer 3 | Unexpected call sequence or schema | Tool registry plus output validation | Low |
| V. Data Exfiltration | Layers 3 + 4 | Egress to an unregistered endpoint | Egress monitor plus PII classifier | Medium |
| VI. Privilege Escalation | Layers 2 + 3 | Permission request from a sub-agent | Permission matrix plus rate limiter | Medium |
The average cost of a single data breach reached $4.88 million in 2024 [1], before counting the costs that attach specifically to agent-mediated incidents, the program suspension, the vendor re-review, the re-approval cycle, and the credibility repair with regulators and customers. A full AASF control stack, by contrast, is a six-to-twelve-week build for a team of four to five practitioners, directionally.
That is the asymmetry. The control stack costs a small, known, one-time fraction of the benchmark figure, while the uncontrolled downside is open-ended and compounding. The curve in Exhibit 5 shows how expected incident cost falls as governance maturity rises, and the heatmap in Exhibit 6 shows exactly where each increment of maturity does its work.
Cost of inaction. An agent-mediated incident carries the benchmark cost [1] plus the exposure unique to autonomy shown in Exhibit 7. During the freeze, every workflow the agent had absorbed returns to manual headcount.
Cost of control. The full stack is a six-to-twelve-week build for four to five practitioners. Estates scale sublinearly, because Layers 1, 2, and 4 are shared infrastructure every additional agent inherits.
Payback logic. One avoided material incident exceeds the full build cost by an order of magnitude. The stack does not need to fire often to justify itself. It needs to fire once.
Each gate is a measurable exit condition, not a status meeting. A program that cannot pass Gate 1 on a single pilot agent has no business scaling to an estate.
Does the agent only read and summarize, or can it send, write, purchase, and delete in external systems? Reach, not model quality, sets the exposure ceiling.
Can every action be rolled back? An email sent or a payment issued cannot. Irreversible reach demands a human gate regardless of model accuracy.
A single agent is auditable end to end. A mesh of delegating agents multiplies Class VI surface with every edge and needs the permission matrix from day one.
Govern to the most sensitive item inside the agent's Recall Boundary, not to the average, because the attacker targets the maximum.
The closing argument. Agent capability is compounding faster than agent governance, and that gap is the largest unpriced risk in enterprise AI today. The organizations that close it now will run autonomy as a competitive advantage. The framework is in these pages. The fifteen weeks start whenever you do.
Answer honestly for your most capable deployed agent. Count the YES answers, then find your level on the band below and your expected exposure on the Exhibit 5 curve. Designed to be completed in pen.
| No. | Probes | Question | Yes | No |
|---|---|---|---|---|
| 1 | L3 | Can you name every tool each deployed agent is able to invoke? | ||
| 2 | L1 | Is every agent input, including retrieved documents, screened before it reaches the model? | ||
| 3 | L2 | Would you detect an agent whose objective had been silently redirected mid-task? | ||
| 4 | L3 | Does every irreversible action pass a human approval gate? | ||
| 5 | L2 | Is there a defined Recall Boundary for each agent, with alerts when it is exceeded? | ||
| 6 | L2 | Could you reconstruct, from logs alone, every action an agent took last Tuesday? | ||
| 7 | L3 | Is agent-to-agent permission inheritance explicitly governed by a matrix? | ||
| 8 | L4 | Does outbound agent traffic pass an egress monitor with PII classification? | ||
| 9 | OPS | Have you red-teamed an agent, with model-driven adversarial probing, in the last two quarters? | ||
| 10 | L2 | Could you roll back everything an agent did in the last hour? |
This paper is a conceptual contribution. The AASF taxonomy, the four-layer control architecture, and the posture model are original to this work. The constructs Action Blast Radius, Memory Drift, and Recall Boundary are formally defined in the author's prior concept papers [7][8] and applied here to the agent security problem. All are grounded in the published incident and adversarial-testing research cited in the notes, and in practitioner observation from enterprise pilot deployments.
Quantitative claims follow one discipline throughout, made visible in every exhibit. Solid fills carry cited data. Hatched fills and curves labeled directional exist to make the structure of an argument legible, not to report measurements. No statistic in this paper is attributed to a source that cannot be independently verified.
Arjun Jaggi writes and advises at the intersection of enterprise AI strategy, governance, and security. His published body of work, spanning original frameworks, concept papers, and books, is read by executive teams navigating AI adoption and is available in full, without a paywall, at arjunjaggi.com.