First Edition · 2026

The Inference
Budget

A Governance Framework for Enterprise AI Inference Cost Accountability

ARJUN JAGGI
ADITYA KARNAM GURURAJ RAO
arjunjaggi.com  ·  adityakarnam.com
THE INFERENCE BUDGET TABLE OF CONTENTS
Front Matter
Executive Summary
ES
Front Matter
Foreword
FW
Chapter 01
The Cost Visibility Gap
01
Chapter 02
Inference Sedimentation
02
Chapter 03
The Routing Imperative
03
Chapter 04
Budget Boundary Failure
04
Chapter 05
The Attribution Problem
05
Chapter 06
The Four Controls
06
Chapter 07
The 60-Day Sprint
07
Chapter 08
The Audit Layer
08
Back Matter
Authors  ·  References
REF
TIB
ES
Executive Summary

Five Findings Every CFO and CTO Must Act On

Enterprise AI inference spend is growing faster than any other IT cost category, yet most organizations lack the governance mechanisms to see, attribute, or control it.

Executive Summary Five Findings Every CFO and CTO Must Act On
Key Findings
  • Finding 1. Most enterprises cannot attribute their AI inference spend to specific decisions, teams, or business outcomes. The root cause is not a data problem. It is an architectural one: inference calls are made without cost metadata, and no governance layer connects spend to purpose.
  • Finding 2. Inference costs do not stay flat between budget reviews. Suboptimal patterns compound over time through a mechanism we term Inference Sedimentation. No single call causes the drift. The accumulation is invisible until it surfaces as a budget overrun.
  • Finding 3. Not every inference request requires frontier model capability. The gap between what an organization routes to frontier models and what those requests actually require is the single most recoverable cost lever available. Research on cascade routing demonstrates this gap is structurally addressable [2].
  • Finding 4. Authorized inference budgets are regularly exceeded through mechanisms outside the budget owner's visibility. We term this Budget Boundary Failure. The failure modes are predictable and preventable. Research on token budget enforcement in local LLM agents provides formal grounding for the enforcement mechanisms [1].
  • Finding 5. A four-control framework closes the governance gap. Visibility, Routing, Enforcement, and Attribution together constitute an Inference Budget program. Published AI governance frameworks do not specify inference spend governance. This book fills that gap.
The Inference Budget Foreword

Foreword

When organizations began deploying large language models at scale, the dominant governance question was about accuracy and safety. Cost was a secondary concern, something the engineering team would optimize once the capability was proven. That ordering has inverted. Today, inference spend is one of the fastest-growing line items in the enterprise IT budget, and the governance frameworks written to manage AI risk say almost nothing about it.

This book is a response to a specific and correctable structural gap. The two most widely adopted AI governance frameworks, enterprise AI governance frameworks and applicable AI governance standards, address model risk, bias, transparency, and accountability. Neither framework specifies how organizations should instrument inference calls, enforce budget boundaries, attribute spend to business decisions, or route requests to the appropriate model tier. This is not a criticism. Inference cost governance did not exist as a discipline when those frameworks were written. It does now.

We coined two terms in this book because the field lacked vocabulary for what we observed. Inference Sedimentation names the compounding accumulation of suboptimal inference patterns that inflate cost over time without any single identifiable cause. Budget Boundary Failure names the mechanism by which actual inference spend systematically exceeds authorized spend through shadow calls, context bloat, and uncontrolled retry cascades. Both terms are designed to pass the test of immediate usability. A practitioner should be able to use them in a budget review meeting without needing to explain them.

The frameworks, controls, and roadmaps in this book are designed to be executable. Every chapter ends with a practitioner section on what to do immediately. A reader who finishes this book should be able to hand it directly to their team and begin implementation. That is the standard we held ourselves to.

01
Chapter 01

The Cost Visibility Gap

Enterprise AI inference spend is largely invisible. Not because the data does not exist, but because no governance layer connects individual calls to cost, team, or business purpose.

Chapter 01 The Cost Visibility Gap
Key Takeaways
  • Visibility is an architectural property, not a reporting one. Instrumentation must be designed in at the call layer, not retrofitted after the fact.
  • The visibility gap is deepest at the context and retry layers, where cost accumulates without any individual decision to attribute.
  • An organization that cannot measure its inference spend cannot govern it. Measurement is the precondition for every other control in this framework.
Questions for Leadership
  1. What fraction of our inference calls today carry enough metadata to be attributed to a specific team, product, or business decision?
  2. Who owns the inference cost instrument layer, and what is their mandate?
  3. How long does it take us to answer the question: which team generated the most inference spend last month?
Visibility Completeness
V(S) = Cvisible / Ctotal V < 1 indicates a structural visibility gap
The Inference Budget Chapter 01 · The Cost Visibility Gap

Why Inference Spend Is Invisible

Enterprise AI systems make inference calls continuously. Each call has a cost determined by model tier, token count, and provider pricing. Yet in most organizations, these calls are made without the metadata needed to answer the most basic governance questions. Which team made the call. What product triggered it. Whether the result was used. The data exists in provider logs, but no governance layer connects it to the organizational context in which it was made.

This is a design gap, not a data gap. The solution is not a better dashboard on top of existing logs. It is a change to how inference calls are issued: every call must carry a cost attribution header that identifies the team, product, and business purpose at the moment of issuance. Retrofitting attribution after the fact is unreliable and operationally expensive.

The Five Visibility Layers

Inference cost accumulates across five distinct layers, each with different visibility characteristics. At the API call layer, spend is directly visible in provider invoices. At the context window layer, cost is proportional to token count but rarely measured per call. At the retry layer, costs from failed or timed-out calls accumulate silently. At the agent loop layer, multi-step chains can generate orders of magnitude more tokens than the original request. At the shadow call layer, calls made by autonomous agents outside of sanctioned workflows may not appear in any organizational cost view at all.

Directional illustration. Not derived from systematic survey data.
Research Finding

BudgetBench (Rao, Jaggi, arXiv:2609.13149) introduces a budget-tiered protocol for evaluating memory strategies in local LLM agents, demonstrating that token budget enforcement requires explicit architectural mechanisms. Without them, agent loops routinely exceed authorized context budgets through incremental drift rather than a single identifiable overage event.

What to Do Monday Morning

Pull last month's inference invoices and ask your engineering lead to attribute each line item to a team and product. If this exercise takes more than a few hours or produces more than twenty percent unknowns, you have a visibility gap that requires architectural remediation, not just better reporting.

Chapter Takeaways

  • Visibility requires instrumentation at the call layer, not post-hoc log analysis.
  • The five cost layers have fundamentally different visibility characteristics. Treat them separately.
  • V(S) below 0.80 indicates a governance-level visibility failure, not an engineering inconvenience.

Before the Next Budget Review

  1. Establish a baseline: what fraction of calls can you attribute today?
  2. Assign ownership of the instrumentation layer to a named team.
  3. Set a six-week target for V(S) improvement and define the measurement method.
02
Chapter 02

Inference Sedimentation

Suboptimal inference patterns accumulate invisibly over time. No single call causes the problem. The compounding does.

Chapter 02 Inference Sedimentation
Key Takeaways
  • Inference Sedimentation is the accumulation of suboptimal inference patterns, oversized context windows, underutilized caching, unrouted model selection, that compound invisibly over time to inflate cost without any single decision causing the drift.
  • Sedimentation is not caused by waste. It is caused by the absence of a governance feedback loop that would surface and correct each pattern individually.
  • The practical test for sedimentation is simple. If your inference cost per unit output has increased over the past two quarters without a corresponding increase in capability delivered, sedimentation is the most likely explanation.
Questions for Leadership
  1. Has our cost-per-useful-inference-call trended up, down, or flat over the past six months?
  2. Do we have a process for retiring inference patterns that were introduced during a pilot and never optimized for production volume?
  3. Who is responsible for detecting sedimentation before it reaches a budget review?
Sedimentation Accumulation
S(T) = Σt=1T δc(t) δc(t) = cost drift at period t, summed over T periods
The Inference Budget Chapter 02 · Inference Sedimentation

How Sedimentation Forms

Inference Sedimentation forms through a predictable sequence. A team deploys an AI feature during a pilot. The context window is set generously to ensure quality. The model tier is set to frontier to avoid accuracy complaints. Caching is disabled because the latency benefit is not needed at pilot scale. The feature ships and the team moves on. The inference pattern is never revisited. At pilot scale, the cost is negligible. At production scale, it is not.

This sequence repeats across every team deploying AI features in an organization. Each individual pattern is defensible in isolation. None of the engineers made a wrong decision. But the aggregate effect, when dozens of features carry the same unoptimized defaults into production, is a cost structure that no single decision created and no single team owns.

The Three Sedimentation Mechanisms

Context bloat is the most common mechanism. System prompts grow as teams add instructions defensively. Document contexts are set to maximum length. Few prompts are ever shortened once written. Over time, the average token count per call drifts upward without any deliberate decision to increase it.

Model tier drift is the second mechanism. The path of least resistance in any organization is to route every request to the most capable available model. Routing logic that would match lower-complexity requests to smaller models requires engineering investment and ongoing calibration. Without explicit governance incentives, it does not happen.

Cache bypass is the third mechanism. Prompt caching can eliminate inference costs for repeated context elements, but it requires that prompts be structured to maximize cache hits. Most prompts are not written with caching in mind, and the optimization is rarely prioritized after initial deployment.

Directional illustration. Not derived from systematic survey data.
What to Do Monday Morning

Identify your three highest-volume inference call patterns. For each one, check the average token count per call over the past 90 days. If the trend is upward without a corresponding increase in quality metrics, you are observing sedimentation in real time. The remediation is to assign an owner to each pattern with a cost-efficiency target and a review cadence.

Chapter Takeaways

  • Sedimentation is a systemic pattern, not individual waste. It requires systemic governance, not individual accountability.
  • The three mechanisms (context bloat, model tier drift, cache bypass) each require a different remediation intervention.
  • S(T) can be computed from invoice data. Track it quarterly as a leading indicator of governance failure.

Before the Next Budget Review

  1. Measure S(T) for the past four quarters and establish a baseline trend line.
  2. Identify the three teams contributing the most to sedimentation growth.
  3. Create a governance feedback loop that surfaces sedimentation to team leads before it reaches the budget review.
03
Chapter 03

The Routing Imperative

Not every inference request requires frontier model capability. The gap between what organizations route to frontier models and what those requests actually require is the most recoverable cost lever available.

Chapter 03 The Routing Imperative
Key Takeaways
  • Cascade routing matches request complexity to model tier, achieving substantial cost reduction without accuracy degradation on appropriately classified requests.
  • FrugalGPT [2] provides formal grounding for the cascade routing approach, demonstrating that intelligent model selection significantly reduces cost while maintaining quality.
  • Routing logic requires ongoing calibration. A routing model trained on last quarter's request distribution may be miscalibrated for this quarter's patterns.
Questions for Leadership
  1. What fraction of our current inference calls are routed to a model tier higher than the task complexity warrants?
  2. Do we have a routing layer, and if so, who owns its calibration?
  3. What is the cost difference between our current routing and an optimally routed baseline?
Routing Efficiency
Reff = (Cunrouted − Crouted) / Cunrouted Reff = fraction of cost recoverable through intelligent routing
The Inference Budget Chapter 03 · The Routing Imperative

What Routing Actually Does

A routing layer sits between the application and the inference provider. When a request arrives, the router classifies its complexity and routes it to the least expensive model tier capable of producing an acceptable result. Simple classification tasks, short-form summarization, and FAQ-style question answering rarely require frontier model capability. Routing these to smaller, less expensive models reduces cost without degrading the quality that users actually experience.

FrugalGPT [2] provides the formal basis for this approach. Chen, Zaharia, and Zou demonstrate that a cascade routing strategy, in which requests are tried against successively more capable models until a quality threshold is met, significantly reduces LLM API costs while maintaining overall output quality. The key insight is that most requests in a realistic enterprise distribution are easier than the hardest requests the organization encounters, and frontier capability is only required for that harder tail.

Building a Routing Layer

An effective routing layer requires three components. First, a complexity classifier that takes an incoming request and outputs a tier assignment. Second, a quality evaluation function that can determine whether a lower-tier model's output meets the threshold required for the use case. Third, an escalation mechanism that promotes requests to the next tier when the current tier fails to meet the quality threshold.

The complexity classifier is the hardest component to calibrate. Request complexity is not well-defined in general, and a classifier trained on one task distribution will generalize poorly to a different one. Organizations that deploy routing layers successfully treat classifier calibration as an ongoing operational task, not a one-time engineering effort.

Directional illustration. Not derived from systematic survey data.
Research Finding

FrugalGPT (Chen, Zaharia, Zou, arXiv:2305.05176) demonstrates that LLM cascade routing can substantially reduce API costs while maintaining output quality. The paper shows that intelligent model selection, rather than defaulting to the most capable model, recovers a significant fraction of inference spend without accuracy degradation.

Chapter Takeaways

  • Routing is the single highest-leverage cost control available to most enterprises today.
  • Cascade routing [2] is the correct architecture. Avoid binary tier assignment in favor of escalation chains.
  • Reff should be measured quarterly. A declining Reff indicates classifier drift requiring recalibration.

Before the Next Budget Review

  1. Sample 200 inference calls from last month and manually classify them by complexity tier.
  2. Calculate the cost difference between actual routing and your estimated optimal routing.
  3. Use this delta to build the business case for a routing layer investment.
04
Chapter 04

Budget Boundary Failure

Authorized inference budgets are exceeded not through deliberate overspending, but through structural mechanisms that fall outside the budget owner's visibility.

Chapter 04 Budget Boundary Failure
Key Takeaways
  • Budget Boundary Failure occurs when actual inference spend systematically exceeds authorized spend through shadow API calls, context bloat from agent loops, uncontrolled retry cascades, or model tier drift, none of which require explicit budget-exceeding decisions.
  • The failure is structural, not behavioral. Enforcement mechanisms must be architectural, not procedural.
  • BudgetBench [1] provides formal grounding for token budget enforcement in LLM agent contexts, establishing that without explicit budget gates, spend overruns are a predictable architectural outcome.
Questions for Leadership
  1. Does our inference budget have hard enforcement gates, or soft alerts?
  2. Can an autonomous agent in our environment make inference calls that exceed its team's monthly budget without triggering an alert?
  3. What is the mean time between a budget boundary being crossed and someone being informed?
Budget Boundary Failure Probability
Fbb = P(actualspend > authorizedspend | no enforcement) Fbb approaches 1 as agent autonomy and volume increase
The Inference Budget Chapter 04 · Budget Boundary Failure

The Five Failure Modes

Budget Boundary Failure occurs through five predictable mechanisms. Shadow API calls are made by autonomous agents or third-party integrations outside of sanctioned call paths, and their costs may not appear in any organizational cost view. Context bloat from agent loops occurs when a multi-step agent accumulates context across iterations, with each step adding tokens that the original budget authorization did not anticipate. Uncontrolled retry cascades arise when a failed inference call triggers automatic retries without a cost ceiling. Model tier drift occurs when routing defaults silently upgrade from an authorized tier to a more expensive one. And finally, concurrent agent proliferation allows multiple agents to consume budget simultaneously against a budget ceiling designed for sequential usage.

Each of these mechanisms is individually preventable. Together, they represent a structural gap between what enterprise budget processes authorize and what enterprise AI systems actually spend. BudgetBench [1] establishes the formal basis for this gap in the context of local LLM agents, demonstrating that token budget enforcement requires explicit architectural mechanisms. Without them, overruns are a predictable outcome of how agent systems are designed, not an exceptional failure of individual actors.

Hard Gates vs. Soft Alerts

Most organizations respond to Budget Boundary Failure by adding monitoring dashboards and alert thresholds. This is insufficient. An alert that fires after a budget is exceeded does not prevent the overrun. It informs someone about a fact that is already expensive to reverse. Hard enforcement gates, mechanisms that block inference calls when a budget ceiling is reached, are the only reliable control against Budget Boundary Failure.

Directional illustration. Not derived from systematic survey data.
Research Finding

BudgetBench (Rao, Jaggi, arXiv:2609.13149) introduces a budget-tiered protocol for memory strategy evaluation in local LLM agents. The paper demonstrates that context window budget constraints require explicit enforcement mechanisms and that memory strategies which appear efficient at low token budgets can generate significant overruns at scale without hard budget gates.

Chapter Takeaways

  • Budget Boundary Failure is architectural. Only architectural controls prevent it reliably.
  • Soft alerts are necessary but not sufficient. Hard budget gates are the mandatory control.
  • Design agent retry logic with explicit cost ceilings, not just attempt count limits.

Before the Next Budget Review

  1. Audit whether any inference call path in your environment lacks a hard cost ceiling.
  2. Identify the top three Budget Boundary Failure modes from the list above that apply to your current architecture.
  3. Assign remediation ownership for each and set a 30-day target for architectural fixes.
05
Chapter 05

Attribution as Infrastructure

Cost attribution is not a finance function. It is an engineering discipline that must be designed into the inference layer, not derived from it after the fact.

Chapter 05 Attribution as Infrastructure
Key Takeaways
  • Attribution completeness A measures what fraction of total inference spend can be traced to a responsible team, product, or decision. A below 0.85 indicates a structural governance failure.
  • Attribution requires a cost header on every inference call, issued at the moment of the call, not reconstructed from logs afterward.
  • Without attribution, budget accountability is impossible. A budget holder cannot be accountable for spend they cannot see broken down by its sources.
Questions for Leadership
  1. Can we produce a breakdown of inference spend by team and product within 24 hours for any given day?
  2. Who owns the attribution standard, and is it enforced at the API gateway level or just by convention?
  3. What is our current attribution completeness score, and what is our target for next quarter?
Attribution Completeness
A = Cattributed / Ctotal A < 0.85 = governance failure requiring remediation
The Inference Budget Chapter 05 · Attribution as Infrastructure

The Attribution Gap

Most enterprises that have deployed AI at scale can produce a monthly total for inference spend. Very few can produce that spend broken down by team, product, use case, and decision context in real time. This gap is not a reporting problem. It is an architectural problem: the inference calls that generate the spend were issued without the metadata needed to attribute them, and no amount of log analysis can reliably reconstruct what was never captured.

The solution is a cost attribution header, a small set of mandatory fields attached to every inference call at the moment of issuance, identifying the team, product, feature, and business purpose behind the call. This is not a new concept. It mirrors the request attribution patterns used in distributed tracing systems. What is new is applying it systematically to inference calls across an enterprise AI stack, where the consequences of missing attribution accumulate in the form of unaccountable spend.

What an Attribution Header Contains

A minimal attribution header requires four fields. Team identifier, naming the organizational unit responsible for the cost. Product identifier, naming the application or feature that triggered the call. Feature identifier, naming the specific capability within the product. Business context, a short classification of the business purpose (customer service, internal tooling, document processing, and so on). These four fields are sufficient to answer the attribution questions that governance requires, without adding significant overhead to the call issuance process.

Directional illustration. Not derived from systematic survey data.
What to Do Monday Morning

Audit your most recent inference invoice. Identify the ten largest cost line items. For each, determine whether you can attribute it to a specific team and product today without manual investigation. Any line item that cannot be attributed in under five minutes represents an attribution gap requiring architectural remediation.

Chapter Takeaways

  • Attribution is an engineering discipline, not a finance function. Design it in.
  • Four fields are sufficient for governance-grade attribution. Complexity is not the barrier to implementation.
  • Enforce attribution at the API gateway, not by convention. Convention compliance degrades over time.

Before the Next Budget Review

  1. Define the mandatory attribution header fields for your organization.
  2. Implement enforcement at the API gateway level, not as a team-by-team convention.
  3. Set a target A score and measure it monthly.
06
Chapter 06

The Four Controls Framework

Inference cost governance requires four controls operating simultaneously. A program missing any one of the four cannot claim to govern its inference spend.

Chapter 06 The Four Controls Framework
Key Takeaways
  • The governance score G(C) is the minimum of the four control scores: Visibility V, Routing R, Enforcement E, and Attribution A. A program scores only as high as its weakest control.
  • Controls are not independent. Enforcement without Attribution assigns penalties to the wrong teams. Routing without Visibility cannot be calibrated. The four controls form a system, not a checklist.
  • The practical goal is not perfection. A G(C) of 0.85 across all four controls represents a governance-grade program for most enterprise contexts.
Questions for Leadership
  1. Have we scored all four controls for our current program? If not, which control have we never formally assessed?
  2. What is our current G(C), and which of the four controls is limiting it?
  3. Do we have a quarterly governance review process, and does it assess all four controls explicitly?
Governance Score
G(C) = min(V, R, E, A) G(C) limited by the weakest of the four controls
The Inference Budget Chapter 06 · The Four Controls Framework

Why Minimum, Not Average

The governance score G(C) uses a minimum function rather than an average because inference cost governance is a system of interdependent controls, not a portfolio of independent measures. A program with perfect visibility, routing, and attribution but no enforcement has no mechanism to prevent spend overruns. A program with enforcement and attribution but poor visibility cannot identify which patterns to enforce against. Each control depends on the others being present and functional. The minimum function captures this dependency: a missing control collapses the governance score regardless of how well the other three are operating.

This framing has a practical implication for program prioritization. When an organization asks where to invest next in inference governance, the answer is always the lowest-scoring control. Improving a control that is already at 0.90 while the weakest control scores 0.40 does not improve the governance score. Only lifting the minimum matters.

Scoring Each Control

Each of the four controls can be scored on a 0-to-1 scale using observable metrics. Visibility V is the fraction of inference calls that can be attributed to a team and product within 24 hours. Routing efficiency R is the fraction of cost saved by routing versus an unrouted baseline. Enforcement E is the fraction of months in which no team exceeded its authorized budget without an explicit override approval. Attribution completeness A is the fraction of total spend that can be traced to a specific team, product, and business purpose.

These metrics are not precisely defined, and organizations will have different measurement capabilities. The value of the framework is not in the precision of the score but in the identification of the weakest control. An organization that cannot score one of the four controls has already identified its most urgent governance gap.

Directional illustration. Not derived from systematic survey data.
Field Observation

Organizations that have implemented all four controls typically find that their limiting control is Enforcement, not Visibility or Attribution. The instrumentation layer often gets built because it is an interesting engineering problem. The routing layer gets built because the cost savings are measurable. Attribution gets built because finance requires it. But enforcement, which requires saying no to teams that want to exceed their budgets, often waits for a budget crisis to create the organizational will to implement it.

Chapter Takeaways

  • Score all four controls before investing in any single one. The minimum determines your governance grade.
  • Controls are interdependent. Build them as a system, not as sequential projects.
  • G(C) of 0.85 is the governance-grade target for enterprise programs.

Before the Next Budget Review

  1. Score each of the four controls for your current program using the metrics defined above.
  2. Identify the minimum. Build the Q3 governance roadmap around lifting it.
  3. Schedule a quarterly governance review that scores all four controls explicitly.
07
Chapter 07

The 90-Day Implementation Sprint

Inference budget governance can reach a functional baseline in 90 days. The sequence matters more than the speed.

Chapter 07 The 90-Day Implementation Sprint
Key Takeaways
  • Phase 1 (days 1-30) establishes measurement. Without a baseline, governance decisions are made on intuition, not evidence.
  • Phase 2 (days 31-60) implements routing and attribution. These two controls are prerequisite to Phase 3.
  • Phase 3 (days 61-90) implements enforcement. Enforcement without Phase 1 and 2 complete generates false positives and organizational friction that undermines the program.
Questions for Leadership
  1. Do we have the organizational mandate to enforce budget limits? If not, Phase 3 will fail regardless of technical readiness.
  2. Who is the executive sponsor for this program, and do they have budget authority over the teams in scope?
  3. What is our definition of success at the end of the 90 days?
Sprint Sequence
Phase 1 Measure Phase 2 Route + Attribute Phase 3 Enforce
The Inference Budget Chapter 07 · The 90-Day Implementation Sprint

Phase 1: Establish Measurement (Days 1-30)

The first 30 days have one deliverable: a baseline measurement of all four control scores using current capabilities. This means pulling every inference invoice from the past 90 days and manually attributing as much spend as possible. It means sampling a set of inference calls and classifying them by complexity to estimate the potential routing efficiency. It means auditing whether any team exceeded its informal budget allocation in the past quarter. And it means estimating current attribution completeness from the metadata that does and does not exist in current logs.

The output of Phase 1 is a governance score card showing the current state of all four controls, with the weakest control identified. This becomes the planning document for Phase 2.

Phase 2: Route and Attribute (Days 31-60)

Phase 2 implements the two controls that are prerequisite to enforcement: routing and attribution. The attribution header standard is defined and implemented at the API gateway. Teams are trained and given a transition period to achieve attribution completeness. The routing architecture is designed and a pilot is deployed covering the highest-volume, most complexity-homogeneous request categories. The routing classifier is calibrated against the Phase 1 sampling data.

Phase 3: Enforce (Days 61-90)

Phase 3 implements budget enforcement using the attribution infrastructure built in Phase 2. Budget allocations are assigned to each team. Hard ceiling mechanisms are deployed, initially in alert-only mode for 30 days to surface false positives and calibrate thresholds. The escalation process for legitimate budget overrides is defined and communicated. At the end of Phase 3, the enforcement controls are switched to hard block mode for teams that have completed the calibration period.

Directional illustration. Not derived from systematic survey data.
What to Do Monday Morning

Before starting the sprint, secure two things. First, a named executive sponsor with budget authority over all teams in scope. Second, a written definition of success: the G(C) target at the end of 90 days, and the four control scores that define it. A sprint without these two foundations will stall in Phase 3 when enforcement decisions require organizational authority to execute.

Chapter Takeaways

  • Sequence matters. Phase 3 fails without Phase 1 and 2 complete.
  • Alert-only mode for 30 days before hard block is the standard approach to prevent false positive friction.
  • Executive sponsorship and a written success definition are prerequisites, not nice-to-haves.

Before the Next Budget Review

  1. Determine whether you have the organizational prerequisites for Phase 3.
  2. If not, do Phase 1 and 2 first and present the Phase 1 results to the executive team to build the mandate for Phase 3.
  3. Assign a program owner for the 90-day sprint before writing a single line of code.
08
Chapter 08

The Inference Audit

Governance without audit is policy without consequence. The inference audit is the mechanism that keeps the four controls honest after the sprint is complete.

Chapter 08 The Inference Audit
Key Takeaways
  • The inference audit is a quarterly process that scores all four controls, identifies new sedimentation patterns, and surfaces emerging Budget Boundary Failure modes.
  • Audit findings require assigned owners and remediation timelines. An audit that produces observations without owners is a compliance exercise, not a governance mechanism.
  • The audit cadence should match the rate of change in the AI program. Organizations adding new AI features faster than quarterly should consider monthly audits.
Questions for Leadership
  1. Do we have a scheduled inference audit process today? If not, who would own it?
  2. What triggers an out-of-cycle audit? A budget overrun? A new AI deployment? A provider pricing change?
  3. How does audit output feed into the annual AI budget planning process?
Audit Cycle
Quarterly Score G(C) Identify Drift Assign Owners Remediate and Retest
The Inference Budget Chapter 08 · The Inference Audit

What the Inference Audit Measures

The quarterly inference audit has four components. First, a re-scoring of all four controls using the same methodology established in Phase 1 of the implementation sprint. This creates a time series of governance scores that surfaces trends, not just snapshots. Second, a sedimentation review that identifies new instances of context bloat, model tier drift, and cache bypass that have accumulated since the last audit. Third, a Budget Boundary Failure review that checks whether any of the five failure modes identified in Chapter 4 have created new overrun patterns. Fourth, a routing calibration review that checks whether the complexity classifier is still accurate for the current request distribution.

Each component produces findings, and each finding must have an assigned owner and a remediation deadline before the audit closes. An audit that produces observations without owners has not produced governance output. It has produced a report that will be filed and forgotten.

Triggering Out-of-Cycle Audits

Four events should trigger an out-of-cycle inference audit. A budget overrun that exceeds the authorized threshold by more than ten percent. A new AI feature deployment that adds more than five percent to the organization's total inference call volume. A material change in provider pricing that affects cost modeling. And any change to the routing architecture, including the introduction of new model tiers or the deprecation of existing ones.

Out-of-cycle audits do not need to cover all four components. A budget overrun triggers a focused Budget Boundary Failure review. A new deployment triggers a sedimentation and attribution review for the new feature. The scope should match the trigger, not default to the full quarterly scope.

Feeding Audit Output into Budget Planning

The single most valuable thing the inference audit can do for an organization is feed its findings into the annual AI budget planning process. An audit that scores G(C) at 0.70 with Enforcement as the limiting control provides a direct input to the budget request: the cost of not improving Enforcement is quantifiable as the fraction of spend that occurs above authorized levels. This transforms the governance program from a cost center into a budget management tool, which is a much stronger organizational position.

Directional illustration. Not derived from systematic survey data.
"Governance without audit is policy without consequence."

Chapter Takeaways

  • The audit produces findings with assigned owners and remediation deadlines, not just observations.
  • Four triggers warrant out-of-cycle audits. Define them in advance and communicate them to teams.
  • Audit output feeds budget planning. This is how governance programs justify their own existence.

Before the Next Budget Review

  1. Schedule the first inference audit for 90 days after the sprint completes.
  2. Define the four out-of-cycle audit triggers and communicate them to engineering leads.
  3. Assign audit ownership now, before the sprint begins, so the owner has context from day one.
Arjun
Jaggi
&
Aditya
Karnam
Gururaj Rao
About the Authors
Arjun Jaggi

Arjun Jaggi is an enterprise AI strategist and framework builder whose work focuses on the governance, safety, and operational deployment of large-scale AI systems. He is the principal researcher behind a body of original frameworks that address structural gaps in enterprise AI governance that no existing standard yet covers. His research combines formal treatment, including mathematical definitions and citable constructs, with immediate practitioner applicability. Organizations and practitioners at the CAO, CISO, and CTO level engage with his work as a primary reference for evaluating AI governance programs against the state of the art.

In this book, Jaggi introduces the formal constructs of Inference Sedimentation and Budget Boundary Failure, identifies the Governance Gap in existing standards, and develops the Four Controls Framework as a complete governance architecture for enterprise inference spend accountability. arjunjaggi.com

Aditya Karnam Gururaj Rao

Aditya Karnam Gururaj Rao is an AI researcher and engineer specializing in evaluation frameworks and model deployment architecture. He is a co-developer of BudgetBench, a budget-tiered protocol for evaluating memory strategies in local LLM agents, published in the IEEE RMKMATE 2025 proceedings. His work spans model evaluation, inference optimization, and the formal measurement of AI system performance across deployment contexts. He has contributed to research on clinical AI deployment and fine-tuning methodologies for specialized domains.

In this book, Rao contributes the formal measurement foundations for the Visibility and Attribution controls, drawing on the BudgetBench methodology to ground the budget enforcement architecture in peer-reviewed evaluation practice.

© 2026 Arjun Jaggi and Aditya Karnam Gururaj Rao. All rights reserved. Academic citation permitted with attribution; commercial use and derivative frameworks require written permission.

The Inference Budget References

References

On Coined Terms: The terms "Inference Sedimentation" and "Budget Boundary Failure" originate with this work and are subject to the copyright notice on the authors page. These terms may be used with attribution in academic and practitioner contexts; commercial use in frameworks, products, or derivative materials requires written permission.