↓ Download PDF
Arjun Jaggi
Enterprise AI Research · White Paper No. 01
THE POINT OF COMPROMISE
October 2026
The Agentic
Attack Surface
A governance framework for securing autonomous AI systems in the enterprise
Author
Arjun Jaggi
Framework
AASF, Version 1.0
Discipline
Enterprise AI Governance
Readership
Boards · C-suite · Security leadership
The Agentic Attack SurfaceArjun Jaggi · 2026
Contents
·
Foreword
A letter to the reader
03
·
Executive Summary
Five findings for the board
04
01
The Problem and the Market Need
Why perimeter security cannot see agentic risk
05
02
The AASF Taxonomy
Six attack classes, one kill chain
06
03
The Control Architecture
Four layers, twenty controls
07
04
The Economics
Cost of inaction and return on control
08
05
The Implementation Roadmap
Fifteen weeks, two hard gates
10
06
The Decision Framework
Four variables, four deployment postures
11
·
The AASF on a Page
The complete framework, one spread, desk-ready
12
·
Self-Assessment Scorecard
Ten questions place your program on the maturity curve
13
·
Methodology, Notes, and the Author
14

About This Paper

This paper introduces the Agentic Attack Surface Framework (AASF), an original taxonomy and control architecture for securing enterprise AI agents. It is written for the leaders who must sign off on agent programs, the CISO who must defend them, the CTO who must build them, and the board that must price the risk of both.

The Agentic Attack Surface Framework, its ten exhibits, the four-layer control architecture, and the posture model are original contributions of this work, published under the license on the back cover. The constructs Action Blast Radius, Memory Drift, and Recall Boundary are drawn from the author's prior research and applied here to agent security, formally defined in the concept papers cited in the notes [7][8]. Every quantitative claim is cited to a named source or labeled directional, without exception.

How to read the exhibits
Solid fill, cited data
Hatched fill, directional illustration
Red, threat path
Arjun Jaggi · Enterprise AI Research02
The Agentic Attack SurfaceForeword
Foreword
A letter to the reader

Every generation of enterprise technology arrives with a security model borrowed from the generation before it, and every generation discovers, at some cost, where the borrowed model breaks. Client-server borrowed from the mainframe. Cloud borrowed from the data center. AI agents are now borrowing from all of them, and the fit is worse than it has ever been.

The reason is simple to state and hard to govern. An agent is the first enterprise system whose inputs are also its instructions. The document it retrieves, the email it reads, the tool output it parses, each one can carry a command, and the agent cannot reliably tell the difference. Researchers demonstrated this against real deployed applications within months of the first agent frameworks shipping [3]. Security teams are being asked to govern a system that takes orders from its own reading material.

This paper is my answer to the question I am asked most often by executive teams. Not whether to deploy agents, that decision has already been made by the market, but how to deploy them so that a single poisoned document cannot become a seven-figure incident. The framework in these pages is original work. The numbers in these pages are cited or labeled directional, without exception. I hold this work to the standard of the rooms it will be read in.

Arjun Jaggi
Enterprise AI Research · arjunjaggi.com

At a Glance

$4.88M
Average cost of a data breach, IBM 2024 [1]
233
Reported AI incidents in 2024, up 56.4% year on year [2]
15 wks
From ad-hoc exposure to governed autonomy, directional
24%
Benchmarked rate at which a tool-using agent obeyed injected instructions, ReAct-prompted frontier model [6]
Arjun Jaggi · Enterprise AI Research03
The Agentic Attack SurfaceExecutive Summary
Executive Summary
Five findings for the board

An AI agent is not a smarter chatbot. It is an autonomous system with persistent memory, live tool access, and the authority to act on behalf of the enterprise. That combination creates an attack surface that perimeter-era security was never designed to see. Five findings follow.

1

Agents collapse the boundary between data and instruction. Every document an agent retrieves and every tool output it reads is a potential command channel. Indirect prompt injection has been demonstrated against real deployed applications [3], not laboratory models. This is a present condition, not a forecast.

2

Exposure scales with tool access, not data footprint. One compromised agent with email, file, and SaaS reach can move data across all three in a single task sequence. Governance must bound the Action Blast Radius, the set of irreversible actions reachable before a human gate, before the program scales.

3

Six attack classes cover the agentic threat space. The AASF taxonomy maps each class to the system layer it enters through and the control layer that intercepts it. A gap in any one layer is reachable from a compromise earlier in the chain, which is why partial stacks fail.

4

Detection probability collapses after goal substitution. Once an agent's objective has been silently redirected, every later action looks legitimate in isolation. Controls inserted before tool invocation return the most protection per engineering hour, a result quantified in Exhibit 6.

5

Governed autonomy is reachable in one quarter. A four-layer control stack deployed across three gated phases takes a single-agent program from ad-hoc exposure to governed autonomy in roughly fifteen weeks, directionally, with a team of four to five practitioners.

Arjun Jaggi · Enterprise AI Research04
The Agentic Attack SurfaceSection 01
01
Section 01 · The Problem
The market has deployed a system its security model cannot see

Enterprise adoption of AI agents has outrun enterprise governance of them, and the gap is now measurable. Reported AI incidents reached 233 in 2024, a 56.4 percent increase over the prior year [2], while the average cost of a data breach reached $4.88 million [1]. Agents raise the stakes on both, multiplying how an incident can begin and what it can reach. The exposure is measured, not hypothetical. In one published benchmark a tool-using frontier model obeyed instructions hidden in retrieved content in 24 percent of 1,054 cases [6].

Traditional controls operate on syntax. Agent threats operate on semantics. A prompt injection attack arrives as a perfectly valid request whose content redirects agent behavior, and no firewall rule, rate limit, or schema check will catch it. The earliest public demonstrations planted instructions in web pages and documents that deployed assistants then obeyed [3][5].

Framework Term

Action Blast Radius. The total set of irreversible external actions an agent can execute before reaching a human approval gate, formalized as a scalar risk metric in the cited concept paper [7]. Every control in this paper exists to shrink it.

EXHIBIT 1
AI incidents are rising at venture speed, not policy speed
Reported AI incidents per year, AI Incident Database
0 100 200 149 233 2023 2024 +56.4%
Source. Stanford HAI, Artificial Intelligence Index Report, 2025 [2].

The structural change is shown in Exhibit 2. An agent sits inside one trust boundary with five distinct system layers, and an adversary can enter through any of them. The damage potential of a compromise is proportional to the agent's tool access, not the data it holds.

EXHIBIT 2
Five system layers sit inside one trust boundary, and each layer is a distinct point of entry
USER INTERFACE ORCHES- TRATION AGENT EXECUTION TOOL LAYER DATA + RETRIEVAL USER / API CLIENT ADMIN CONSOLE EVENTS / WEBHOOKS INGRESS GATEWAY AUTHN + SESSION TASK ROUTER DISPATCH LOGIC CONTEXT MANAGER WINDOW BUDGET POLICY GUARD RULE ENGINE LLM CORE REASONING ENGINE MEMORY SHORT + LONG TERM PLANNER TASK DECOMPOSITION TOOL CALLER FUNCTION DISPATCH WEB CODE EMAIL SAAS FILES APIS VECTOR DB KNOWLEDGE SESSIONS AUDIT LOG EMBEDDINGS ATTACK VECTORS I PROMPT INJECTION via any untrusted input II GOAL HIJACKING via planning manipulation III MEMORY POISONING via persistent stores IV TOOL POISONING via compromised tooling V DATA EXFILTRATION via legitimate egress Highest-value target System component Trust boundary Class VI, privilege escalation, spans lanes two to four
Source. AASF framework, original to this work. The LLM core is the highest-value target because every other layer trusts its output.
Arjun Jaggi · Enterprise AI Research05
The Agentic Attack SurfaceSection 02
02
Section 02 · The Taxonomy
Six attack classes cover the agentic threat space

The Agentic Attack Surface Framework classifies threats by the system layer they enter through and the type of agency they exploit. Every agentic incident observed in practice resolves to one of these six classes, or to a chain of them.

I
Prompt Injection

Malicious instructions embedded in inputs, retrieved documents, or tool outputs redirect agent behavior without the orchestrator's knowledge. Demonstrated against deployed applications [3].

II
Goal Hijacking

Manipulation of multi-step planning so the agent pursues an attacker objective. Hard to detect because each step looks legitimate in isolation.

III
Memory Poisoning

Adversarial content written into persistent or working memory shapes all future reasoning, within and across sessions. The slowest attack and the hardest to unwind.

IV
Tool Poisoning

A compromised tool definition, output schema, or upstream API causes the agent to take unintended actions while believing it executed correctly.

V
Data Exfiltration

The agent's legitimate tool access routes sensitive data to attacker endpoints. The agent acts as a compliant carrier, so no system compromise is required.

VI
Privilege Escalation

Agent-to-agent communication or tool chaining acquires permissions beyond those granted at session start. The signature risk of multi-agent architectures.

EXHIBIT 3
Detection probability collapses after Stage 3, so the cheapest controls sit earliest in the chain
INPUT VALIDATION CONTEXT MONITOR TOOL AUTHORIZATION GATE EGRESS MONITOR 1 2 3 4 5 6 INJECTION Payload enters the context window ACCEPTANCE Payload treated as trusted instruction GOAL SHIFT Objective silently substituted TOOL INVOKE Attacker actions run with agent authority EXFILTRATION Sensitive content moves outbound IMPACT Breach, regulatory exposure, ransom HIGH DETECTION PROBABILITY MODERATE DETECTION POST-INCIDENT FORENSICS ONLY MEAN TIME TO DETECT LENGTHENS AS THE CHAIN ADVANCES
Source. AASF framework, original to this work. The four blue markers are control insertion points; the two left of Stage 3 cost the least and recover the most.
Framework Term

Recall Boundary. The set of information an agent may retrieve and act on within one task session, formally defined in the cited concept paper [8]. Exceeding it without authorization is the earliest reliable Class III signal.

Framework Term

Memory Drift. Gradual divergence between an agent's stored behavioral priors and its operational context, formally defined in the cited concept paper [8]. The gap drift opens is where Class III attacks live.

Arjun Jaggi · Enterprise AI Research06
The Agentic Attack SurfaceSection 03
03
Section 03 · The Architecture
Four control layers, because no single layer survives contact

Each layer intercepts a different stage of the Exhibit 3 chain, and a gap in any layer is reachable from a compromise earlier in the chain. Traffic that clears Layer 1 is still watched by Layer 2, still scoped by Layer 3, and still bounded by Layer 4.

EXHIBIT 4
The AASF defense stack places twenty controls at four moments in the agent lifecycle
1 INPUT VALIDATION AND SANITIZATION PRE-EXECUTION PROMPT CLASSIFIER INJECTION DETECTOR SCHEMA VALIDATOR INTENT SCORER CONTENT FILTER COVERS CLASS I AND II 2 RUNTIME BEHAVIORAL MONITORING IN-EXECUTION ACTION AUDIT LOG GOAL DRIFT MONITOR LOOP DETECTOR BLAST RADIUS LIMIT ROLLBACK ENGINE COVERS CLASS II, III, VI 3 TOOL AUTHORIZATION AND SCOPING AT-ACTION TOOL REGISTRY PERMISSION MATRIX RATE LIMITER OUTPUT VALIDATOR HUMAN GATE COVERS CLASS IV AND V 4 DATA GOVERNANCE AND EGRESS CONTROL AT-REST + OUTBOUND PII CLASSIFIER EGRESS MONITOR REDACTION ENGINE ACCESS AUDIT RETENTION POLICY COVERS CLASS III AND V
Source. AASF framework, original to this work. Layers 1 and 2 intercept before and during reasoning; Layer 3 gates each action at invocation; Layer 4 bounds what any successful compromise can carry out.
Attack classPrimary layerEarliest detection signalMitigationResidual risk
I. Prompt InjectionLayer 1Anomalous instruction pattern in inputClassifier plus schema enforcementMedium
II. Goal HijackingLayers 1 + 2Divergence from declared objectiveGoal drift monitor plus human gateMedium
III. Memory PoisoningLayers 2 + 4Recall Boundary exceededMemory scoping plus access auditHigh
IV. Tool PoisoningLayer 3Unexpected call sequence or schemaTool registry plus output validationLow
V. Data ExfiltrationLayers 3 + 4Egress to an unregistered endpointEgress monitor plus PII classifierMedium
VI. Privilege EscalationLayers 2 + 3Permission request from a sub-agentPermission matrix plus rate limiterMedium
Arjun Jaggi · Enterprise AI Research07
The Agentic Attack SurfaceSection 04
04
Section 04 · The Economics
The asymmetry that makes this decision easy

The average cost of a single data breach reached $4.88 million in 2024 [1], before counting the costs that attach specifically to agent-mediated incidents, the program suspension, the vendor re-review, the re-approval cycle, and the credibility repair with regulators and customers. A full AASF control stack, by contrast, is a six-to-twelve-week build for a team of four to five practitioners, directionally.

That is the asymmetry. The control stack costs a small, known, one-time fraction of the benchmark figure, while the uncontrolled downside is open-ended and compounding. The curve in Exhibit 5 shows how expected incident cost falls as governance maturity rises, and the heatmap in Exhibit 6 shows exactly where each increment of maturity does its work.

$4.88M
The global average total cost of a data breach in 2024, the benchmark every agent program should price its governance against [1]
EXHIBIT 5
Expected incident cost falls by an order of magnitude across five levels of governance maturity
$5M $4M $3M $2M $1M $0 $4.9M $3.6M $2.4M $1.3M $0.6M 1 AD-HOC 2 AWARE 3 MANAGED 4 OPTIMIZED 5 CONTINUOUS GOVERNANCE MATURITY LEVEL
Source. Directional illustration, original to this work, anchored at Level 1 to the IBM 2024 global average breach cost [1]. The steepest savings sit between Levels 1 and 3, the ground covered by Phases 1 and 2 of the Section 05 roadmap.
Arjun Jaggi · Enterprise AI Research08
The Agentic Attack SurfaceSection 04 · Continued
EXHIBIT 6
No attack class is fully covered by any single layer, which is the quantitative case for the full stack
Control effectiveness by attack class and layer, directional, in percent
LAYER 1 LAYER 2 LAYER 3 LAYER 4 I. Prompt Injection II. Goal Hijacking III. Memory Poisoning IV. Tool Poisoning V. Data Exfiltration VI. Privilege Escalation 90 55 30 10 75 80 40 20 30 65 50 70 20 40 95 30 15 50 85 90 20 75 70 40 LOW HIGH
Source. Directional illustration from pilot deployment observation, original to this work. Class III is the weakest column-maximum in the matrix, consistent with its High residual-risk rating in Section 03.
EXHIBIT 7
An agent incident costs more than the breach itself
Illustrative composition of total exposure, not to scale
BREACH BASE [1] PROGRAM FREEZE RE-APPROVAL CYCLE CREDIBILITY REPAIR TOTAL EXPOSURE CITED HATCHED SEGMENTS ARE DIRECTIONAL
Source. Base figure IBM 2024 [1]; increments are directional, from practitioner observation of post-incident program freezes.

The Return on Control

Cost of inaction. An agent-mediated incident carries the benchmark cost [1] plus the exposure unique to autonomy shown in Exhibit 7. During the freeze, every workflow the agent had absorbed returns to manual headcount.

Cost of control. The full stack is a six-to-twelve-week build for four to five practitioners. Estates scale sublinearly, because Layers 1, 2, and 4 are shared infrastructure every additional agent inherits.

Payback logic. One avoided material incident exceeds the full build cost by an order of magnitude. The stack does not need to fire often to justify itself. It needs to fire once.

Arjun Jaggi · Enterprise AI Research09
The Agentic Attack SurfaceSection 05
05
Section 05 · The Roadmap
Fifteen weeks, three phases, two hard gates

Each gate is a measurable exit condition, not a status meeting. A program that cannot pass Gate 1 on a single pilot agent has no business scaling to an estate.

EXHIBIT 8
The deployment timeline front-loads the controls that Exhibit 3 shows recover the most
WK 1-2 WK 3-4 WK 5-6 WK 7-10 WK 11-14 WK 15+ PHASE 1 Pilot PHASE 2 Harden PHASE 3 Scale Threat model + attack surface map Layer 1 controls Prompt classifier Pilot agent + behavioral baseline GATE 1 Layer 2 behavioral monitoring Tool authorization Egress monitoring Red team exercise + tabletop GATE 2 All-layer controls, estate-wide Continuous monitoring + quarterly red team Phase 1 Phase 2 Phase 3 Go / no-go gate
Source. AASF deployment model, original to this work. Phase 2 includes a structured red team exercise; model-driven adversarial probing scales that exercise well beyond what a human-only team covers [4].
1
Weeks 1-6
Pilot and threat model
  • Select one agent deployment as the pilot
  • Build the complete AASF attack surface map
  • Deploy Layer 1 input validation controls
  • Establish the behavioral audit baseline
  • Gate 1. Two-week run, zero Class I or II incidents
2
Weeks 7-14
Hardening and coverage
  • Deploy Layers 2 through 4
  • Implement the tool permission matrix
  • Run the first structured red team exercise
  • Instrument every agent-to-agent call
  • Gate 2. Zero unmitigated Class V or VI paths
3
Weeks 15+
Enterprise rollout
  • Extend the framework to every agent in the estate
  • Automate control validation in CI/CD
  • Set a quarterly red team cadence
  • Publish the internal agent security policy
  • Success. Classifier false positives under one percent
Arjun Jaggi · Enterprise AI Research10
The Agentic Attack SurfaceSection 06
06
Section 06 · The Decision
Four variables select the posture, and the posture sets the stack
1
Action Reach

Does the agent only read and summarize, or can it send, write, purchase, and delete in external systems? Reach, not model quality, sets the exposure ceiling.

2
Reversibility

Can every action be rolled back? An email sent or a payment issued cannot. Irreversible reach demands a human gate regardless of model accuracy.

3
Agent Topology

A single agent is auditable end to end. A mesh of delegating agents multiplies Class VI surface with every edge and needs the permission matrix from day one.

4
Recall Sensitivity

Govern to the most sensitive item inside the agent's Recall Boundary, not to the average, because the attacker targets the maximum.

EXHIBIT 9
Three questions resolve any agent use case to one of four deployment postures
THE PATH READS LEFT TO RIGHT. THE POSTURE CHOSEN HERE SETS THE PHASE 1 SCOPE IN SECTION 05. AGENT USE CASE EXTERNAL ACTIONS? NO YES IRREVERSIBLE REACH? NO YES MULTI-AGENT MESH? NO YES POSTURE A Read-only assist Layers 1 and 2 only POSTURE B Scoped autonomy Layers 1 to 3 POSTURE C Gated autonomy Full stack + human gate POSTURE D Federated mesh Full stack + permission matrix
Source. AASF posture model, original to this work. An estate runs mixed postures by design; a Posture A research agent and a Posture C procurement agent can share Layer 1 and 2 infrastructure.
EXHIBIT 10
What each posture carries, at a glance
POSTURE L1 L2 L3 L4 ADDED CONTROL TYPICAL USE A · Read-only assist None required Research, summarization B · Scoped autonomy Tool scoping + rate limits Drafting, data pulls C · Gated autonomy Human approval gate Payments, outbound email D · Federated mesh Permission matrix + sub-agent audit Multi-agent operations
Source. AASF posture model, original to this work. Filled circles mark required control layers.

The closing argument. Agent capability is compounding faster than agent governance, and that gap is the largest unpriced risk in enterprise AI today. The organizations that close it now will run autonomy as a competitive advantage. The framework is in these pages. The fifteen weeks start whenever you do.

Arjun Jaggi · Enterprise AI Research11
The Agentic Attack SurfaceThe Framework
Pin This Page
The AASF on a page
The complete framework in one spread. Everything here is unpacked in Sections 01 to 06.
01 · Six attack classes
I
Prompt Injection
II
Goal Hijacking
III
Memory Poisoning
IV
Tool Poisoning
V
Data Exfiltration
VI
Privilege Escalation
02 · One kill chain, four interception points
INJECTION ACCEPTANCE GOAL SHIFT TOOL INVOKE EXFILTRATION IMPACT DETECTION PROBABILITY FALLS LEFT TO RIGHT. BLUE TICKS MARK CONTROL INSERTION POINTS.
03 · Four control layers
1INPUT VALIDATION AND SANITIZATIONCOVERS I AND II
2RUNTIME BEHAVIORAL MONITORINGCOVERS II, III, VI
3TOOL AUTHORIZATION AND SCOPINGCOVERS IV AND V
4DATA GOVERNANCE AND EGRESS CONTROLCOVERS III AND V
04 · Four deployment postures
POSTURE A
Read-only assist
Layers 1 and 2 only
POSTURE B
Scoped autonomy
Layers 1 to 3
POSTURE C
Gated autonomy
Full stack + human gate
POSTURE D
Federated mesh
Full stack + permission matrix
05 · Fifteen weeks, two hard gates
PHASE 1 · PILOT · WK 1-6 PHASE 2 · HARDEN · WK 7-14 PHASE 3 · SCALE · WK 15+ GATE 1 GATE 2
Action Blast Radius Memory Drift Recall Boundary
One avoided incident pays for the entire stack.
Arjun Jaggi · Enterprise AI Research12
The Agentic Attack SurfaceSelf-Assessment
Self-Assessment · Ten Questions
Where does your program stand?

Answer honestly for your most capable deployed agent. Count the YES answers, then find your level on the band below and your expected exposure on the Exhibit 5 curve. Designed to be completed in pen.

No.ProbesQuestionYesNo
1L3Can you name every tool each deployed agent is able to invoke?
2L1Is every agent input, including retrieved documents, screened before it reaches the model?
3L2Would you detect an agent whose objective had been silently redirected mid-task?
4L3Does every irreversible action pass a human approval gate?
5L2Is there a defined Recall Boundary for each agent, with alerts when it is exceeded?
6L2Could you reconstruct, from logs alone, every action an agent took last Tuesday?
7L3Is agent-to-agent permission inheritance explicitly governed by a matrix?
8L4Does outbound agent traffic pass an egress monitor with PII classification?
9OPSHave you red-teamed an agent, with model-driven adversarial probing, in the last two quarters?
10L2Could you roll back everything an agent did in the last hour?
Your score, your maturity level
0-2 YES
LEVEL 1 · AD-HOC
3-4 YES
LEVEL 2 · AWARE
5-6 YES
LEVEL 3 · MANAGED
7-8 YES
LEVEL 4 · OPTIMIZED
9-10 YES
LEVEL 5 · CONTINUOUS
Scoring bands map one-to-one onto the maturity levels of Exhibit 5. Your level on that curve is the directional price of standing still.
Arjun Jaggi · Enterprise AI Research13
The Agentic Attack SurfaceMethodology · Notes · Author

About the Research

This paper is a conceptual contribution. The AASF taxonomy, the four-layer control architecture, and the posture model are original to this work. The constructs Action Blast Radius, Memory Drift, and Recall Boundary are formally defined in the author's prior concept papers [7][8] and applied here to the agent security problem. All are grounded in the published incident and adversarial-testing research cited in the notes, and in practitioner observation from enterprise pilot deployments.

Quantitative claims follow one discipline throughout, made visible in every exhibit. Solid fills carry cited data. Hatched fills and curves labeled directional exist to make the structure of an argument legible, not to report measurements. No statistic in this paper is attributed to a source that cannot be independently verified.

Notes

About the Author

AJ
Arjun Jaggi

Arjun Jaggi writes and advises at the intersection of enterprise AI strategy, governance, and security. His published body of work, spanning original frameworks, concept papers, and books, is read by executive teams navigating AI adoption and is available in full, without a paywall, at arjunjaggi.com.

Engage · arjunjaggi.com · calendly.com/arjunjaggi
Arjun Jaggi · Enterprise AI Research14
Arjun Jaggi
Enterprise AI Research
Agent capability is compounding faster than agent governance.
This framework closes the gap.
AASF Framework Action Blast Radius Memory Drift Recall Boundary Four Postures
© 2026 Arjun Jaggi. Original framework. Academic citation permitted with attribution; commercial use requires written permission.
White Paper No. 01 · arjunjaggi.com