A Governance Framework for Enterprise AI Inference Cost Accountability
Enterprise AI inference spend is growing faster than any other IT cost category, yet most organizations lack the governance mechanisms to see, attribute, or control it.
When organizations began deploying large language models at scale, the dominant governance question was about accuracy and safety. Cost was a secondary concern, something the engineering team would optimize once the capability was proven. That ordering has inverted. Today, inference spend is one of the fastest-growing line items in the enterprise IT budget, and the governance frameworks written to manage AI risk say almost nothing about it.
This book is a response to a specific and correctable structural gap. The two most widely adopted AI governance frameworks, enterprise AI governance frameworks and applicable AI governance standards, address model risk, bias, transparency, and accountability. Neither framework specifies how organizations should instrument inference calls, enforce budget boundaries, attribute spend to business decisions, or route requests to the appropriate model tier. This is not a criticism. Inference cost governance did not exist as a discipline when those frameworks were written. It does now.
We coined two terms in this book because the field lacked vocabulary for what we observed. Inference Sedimentation names the compounding accumulation of suboptimal inference patterns that inflate cost over time without any single identifiable cause. Budget Boundary Failure names the mechanism by which actual inference spend systematically exceeds authorized spend through shadow calls, context bloat, and uncontrolled retry cascades. Both terms are designed to pass the test of immediate usability. A practitioner should be able to use them in a budget review meeting without needing to explain them.
The frameworks, controls, and roadmaps in this book are designed to be executable. Every chapter ends with a practitioner section on what to do immediately. A reader who finishes this book should be able to hand it directly to their team and begin implementation. That is the standard we held ourselves to.
Enterprise AI inference spend is largely invisible. Not because the data does not exist, but because no governance layer connects individual calls to cost, team, or business purpose.
Enterprise AI systems make inference calls continuously. Each call has a cost determined by model tier, token count, and provider pricing. Yet in most organizations, these calls are made without the metadata needed to answer the most basic governance questions. Which team made the call. What product triggered it. Whether the result was used. The data exists in provider logs, but no governance layer connects it to the organizational context in which it was made.
This is a design gap, not a data gap. The solution is not a better dashboard on top of existing logs. It is a change to how inference calls are issued: every call must carry a cost attribution header that identifies the team, product, and business purpose at the moment of issuance. Retrofitting attribution after the fact is unreliable and operationally expensive.
Inference cost accumulates across five distinct layers, each with different visibility characteristics. At the API call layer, spend is directly visible in provider invoices. At the context window layer, cost is proportional to token count but rarely measured per call. At the retry layer, costs from failed or timed-out calls accumulate silently. At the agent loop layer, multi-step chains can generate orders of magnitude more tokens than the original request. At the shadow call layer, calls made by autonomous agents outside of sanctioned workflows may not appear in any organizational cost view at all.
BudgetBench (Rao, Jaggi, arXiv:2609.13149) introduces a budget-tiered protocol for evaluating memory strategies in local LLM agents, demonstrating that token budget enforcement requires explicit architectural mechanisms. Without them, agent loops routinely exceed authorized context budgets through incremental drift rather than a single identifiable overage event.
Pull last month's inference invoices and ask your engineering lead to attribute each line item to a team and product. If this exercise takes more than a few hours or produces more than twenty percent unknowns, you have a visibility gap that requires architectural remediation, not just better reporting.
Suboptimal inference patterns accumulate invisibly over time. No single call causes the problem. The compounding does.
Inference Sedimentation forms through a predictable sequence. A team deploys an AI feature during a pilot. The context window is set generously to ensure quality. The model tier is set to frontier to avoid accuracy complaints. Caching is disabled because the latency benefit is not needed at pilot scale. The feature ships and the team moves on. The inference pattern is never revisited. At pilot scale, the cost is negligible. At production scale, it is not.
This sequence repeats across every team deploying AI features in an organization. Each individual pattern is defensible in isolation. None of the engineers made a wrong decision. But the aggregate effect, when dozens of features carry the same unoptimized defaults into production, is a cost structure that no single decision created and no single team owns.
Context bloat is the most common mechanism. System prompts grow as teams add instructions defensively. Document contexts are set to maximum length. Few prompts are ever shortened once written. Over time, the average token count per call drifts upward without any deliberate decision to increase it.
Model tier drift is the second mechanism. The path of least resistance in any organization is to route every request to the most capable available model. Routing logic that would match lower-complexity requests to smaller models requires engineering investment and ongoing calibration. Without explicit governance incentives, it does not happen.
Cache bypass is the third mechanism. Prompt caching can eliminate inference costs for repeated context elements, but it requires that prompts be structured to maximize cache hits. Most prompts are not written with caching in mind, and the optimization is rarely prioritized after initial deployment.
Identify your three highest-volume inference call patterns. For each one, check the average token count per call over the past 90 days. If the trend is upward without a corresponding increase in quality metrics, you are observing sedimentation in real time. The remediation is to assign an owner to each pattern with a cost-efficiency target and a review cadence.
Not every inference request requires frontier model capability. The gap between what organizations route to frontier models and what those requests actually require is the most recoverable cost lever available.
A routing layer sits between the application and the inference provider. When a request arrives, the router classifies its complexity and routes it to the least expensive model tier capable of producing an acceptable result. Simple classification tasks, short-form summarization, and FAQ-style question answering rarely require frontier model capability. Routing these to smaller, less expensive models reduces cost without degrading the quality that users actually experience.
FrugalGPT [2] provides the formal basis for this approach. Chen, Zaharia, and Zou demonstrate that a cascade routing strategy, in which requests are tried against successively more capable models until a quality threshold is met, significantly reduces LLM API costs while maintaining overall output quality. The key insight is that most requests in a realistic enterprise distribution are easier than the hardest requests the organization encounters, and frontier capability is only required for that harder tail.
An effective routing layer requires three components. First, a complexity classifier that takes an incoming request and outputs a tier assignment. Second, a quality evaluation function that can determine whether a lower-tier model's output meets the threshold required for the use case. Third, an escalation mechanism that promotes requests to the next tier when the current tier fails to meet the quality threshold.
The complexity classifier is the hardest component to calibrate. Request complexity is not well-defined in general, and a classifier trained on one task distribution will generalize poorly to a different one. Organizations that deploy routing layers successfully treat classifier calibration as an ongoing operational task, not a one-time engineering effort.
FrugalGPT (Chen, Zaharia, Zou, arXiv:2305.05176) demonstrates that LLM cascade routing can substantially reduce API costs while maintaining output quality. The paper shows that intelligent model selection, rather than defaulting to the most capable model, recovers a significant fraction of inference spend without accuracy degradation.
Authorized inference budgets are exceeded not through deliberate overspending, but through structural mechanisms that fall outside the budget owner's visibility.
Budget Boundary Failure occurs through five predictable mechanisms. Shadow API calls are made by autonomous agents or third-party integrations outside of sanctioned call paths, and their costs may not appear in any organizational cost view. Context bloat from agent loops occurs when a multi-step agent accumulates context across iterations, with each step adding tokens that the original budget authorization did not anticipate. Uncontrolled retry cascades arise when a failed inference call triggers automatic retries without a cost ceiling. Model tier drift occurs when routing defaults silently upgrade from an authorized tier to a more expensive one. And finally, concurrent agent proliferation allows multiple agents to consume budget simultaneously against a budget ceiling designed for sequential usage.
Each of these mechanisms is individually preventable. Together, they represent a structural gap between what enterprise budget processes authorize and what enterprise AI systems actually spend. BudgetBench [1] establishes the formal basis for this gap in the context of local LLM agents, demonstrating that token budget enforcement requires explicit architectural mechanisms. Without them, overruns are a predictable outcome of how agent systems are designed, not an exceptional failure of individual actors.
Most organizations respond to Budget Boundary Failure by adding monitoring dashboards and alert thresholds. This is insufficient. An alert that fires after a budget is exceeded does not prevent the overrun. It informs someone about a fact that is already expensive to reverse. Hard enforcement gates, mechanisms that block inference calls when a budget ceiling is reached, are the only reliable control against Budget Boundary Failure.
BudgetBench (Rao, Jaggi, arXiv:2609.13149) introduces a budget-tiered protocol for memory strategy evaluation in local LLM agents. The paper demonstrates that context window budget constraints require explicit enforcement mechanisms and that memory strategies which appear efficient at low token budgets can generate significant overruns at scale without hard budget gates.
Cost attribution is not a finance function. It is an engineering discipline that must be designed into the inference layer, not derived from it after the fact.
Most enterprises that have deployed AI at scale can produce a monthly total for inference spend. Very few can produce that spend broken down by team, product, use case, and decision context in real time. This gap is not a reporting problem. It is an architectural problem: the inference calls that generate the spend were issued without the metadata needed to attribute them, and no amount of log analysis can reliably reconstruct what was never captured.
The solution is a cost attribution header, a small set of mandatory fields attached to every inference call at the moment of issuance, identifying the team, product, feature, and business purpose behind the call. This is not a new concept. It mirrors the request attribution patterns used in distributed tracing systems. What is new is applying it systematically to inference calls across an enterprise AI stack, where the consequences of missing attribution accumulate in the form of unaccountable spend.
A minimal attribution header requires four fields. Team identifier, naming the organizational unit responsible for the cost. Product identifier, naming the application or feature that triggered the call. Feature identifier, naming the specific capability within the product. Business context, a short classification of the business purpose (customer service, internal tooling, document processing, and so on). These four fields are sufficient to answer the attribution questions that governance requires, without adding significant overhead to the call issuance process.
Audit your most recent inference invoice. Identify the ten largest cost line items. For each, determine whether you can attribute it to a specific team and product today without manual investigation. Any line item that cannot be attributed in under five minutes represents an attribution gap requiring architectural remediation.
Inference cost governance requires four controls operating simultaneously. A program missing any one of the four cannot claim to govern its inference spend.
The governance score G(C) uses a minimum function rather than an average because inference cost governance is a system of interdependent controls, not a portfolio of independent measures. A program with perfect visibility, routing, and attribution but no enforcement has no mechanism to prevent spend overruns. A program with enforcement and attribution but poor visibility cannot identify which patterns to enforce against. Each control depends on the others being present and functional. The minimum function captures this dependency: a missing control collapses the governance score regardless of how well the other three are operating.
This framing has a practical implication for program prioritization. When an organization asks where to invest next in inference governance, the answer is always the lowest-scoring control. Improving a control that is already at 0.90 while the weakest control scores 0.40 does not improve the governance score. Only lifting the minimum matters.
Each of the four controls can be scored on a 0-to-1 scale using observable metrics. Visibility V is the fraction of inference calls that can be attributed to a team and product within 24 hours. Routing efficiency R is the fraction of cost saved by routing versus an unrouted baseline. Enforcement E is the fraction of months in which no team exceeded its authorized budget without an explicit override approval. Attribution completeness A is the fraction of total spend that can be traced to a specific team, product, and business purpose.
These metrics are not precisely defined, and organizations will have different measurement capabilities. The value of the framework is not in the precision of the score but in the identification of the weakest control. An organization that cannot score one of the four controls has already identified its most urgent governance gap.
Organizations that have implemented all four controls typically find that their limiting control is Enforcement, not Visibility or Attribution. The instrumentation layer often gets built because it is an interesting engineering problem. The routing layer gets built because the cost savings are measurable. Attribution gets built because finance requires it. But enforcement, which requires saying no to teams that want to exceed their budgets, often waits for a budget crisis to create the organizational will to implement it.
Inference budget governance can reach a functional baseline in 90 days. The sequence matters more than the speed.
The first 30 days have one deliverable: a baseline measurement of all four control scores using current capabilities. This means pulling every inference invoice from the past 90 days and manually attributing as much spend as possible. It means sampling a set of inference calls and classifying them by complexity to estimate the potential routing efficiency. It means auditing whether any team exceeded its informal budget allocation in the past quarter. And it means estimating current attribution completeness from the metadata that does and does not exist in current logs.
The output of Phase 1 is a governance score card showing the current state of all four controls, with the weakest control identified. This becomes the planning document for Phase 2.
Phase 2 implements the two controls that are prerequisite to enforcement: routing and attribution. The attribution header standard is defined and implemented at the API gateway. Teams are trained and given a transition period to achieve attribution completeness. The routing architecture is designed and a pilot is deployed covering the highest-volume, most complexity-homogeneous request categories. The routing classifier is calibrated against the Phase 1 sampling data.
Phase 3 implements budget enforcement using the attribution infrastructure built in Phase 2. Budget allocations are assigned to each team. Hard ceiling mechanisms are deployed, initially in alert-only mode for 30 days to surface false positives and calibrate thresholds. The escalation process for legitimate budget overrides is defined and communicated. At the end of Phase 3, the enforcement controls are switched to hard block mode for teams that have completed the calibration period.
Before starting the sprint, secure two things. First, a named executive sponsor with budget authority over all teams in scope. Second, a written definition of success: the G(C) target at the end of 90 days, and the four control scores that define it. A sprint without these two foundations will stall in Phase 3 when enforcement decisions require organizational authority to execute.
Governance without audit is policy without consequence. The inference audit is the mechanism that keeps the four controls honest after the sprint is complete.
The quarterly inference audit has four components. First, a re-scoring of all four controls using the same methodology established in Phase 1 of the implementation sprint. This creates a time series of governance scores that surfaces trends, not just snapshots. Second, a sedimentation review that identifies new instances of context bloat, model tier drift, and cache bypass that have accumulated since the last audit. Third, a Budget Boundary Failure review that checks whether any of the five failure modes identified in Chapter 4 have created new overrun patterns. Fourth, a routing calibration review that checks whether the complexity classifier is still accurate for the current request distribution.
Each component produces findings, and each finding must have an assigned owner and a remediation deadline before the audit closes. An audit that produces observations without owners has not produced governance output. It has produced a report that will be filed and forgotten.
Four events should trigger an out-of-cycle inference audit. A budget overrun that exceeds the authorized threshold by more than ten percent. A new AI feature deployment that adds more than five percent to the organization's total inference call volume. A material change in provider pricing that affects cost modeling. And any change to the routing architecture, including the introduction of new model tiers or the deprecation of existing ones.
Out-of-cycle audits do not need to cover all four components. A budget overrun triggers a focused Budget Boundary Failure review. A new deployment triggers a sedimentation and attribution review for the new feature. The scope should match the trigger, not default to the full quarterly scope.
The single most valuable thing the inference audit can do for an organization is feed its findings into the annual AI budget planning process. An audit that scores G(C) at 0.70 with Enforcement as the limiting control provides a direct input to the budget request: the cost of not improving Enforcement is quantifiable as the fraction of spend that occurs above authorized levels. This transforms the governance program from a cost center into a budget management tool, which is a much stronger organizational position.
On Coined Terms: The terms "Inference Sedimentation" and "Budget Boundary Failure" originate with this work and are subject to the copyright notice on the authors page. These terms may be used with attribution in academic and practitioner contexts; commercial use in frameworks, products, or derivative materials requires written permission.