The Budget Conversation That Ends Programs
A CIO presents their AI program results at the quarterly business review. They show query volumes, average response time, user adoption rates, and estimated hours saved based on a self-reported survey. The CFO asks one question. What did it do for the business? The CIO restates the metrics. The CFO schedules a follow-up to discuss budget reallocation.
This is not a communication failure. It is a measurement architecture failure. The program was designed, from the beginning, to measure at the wrong level. The CIO has activity data. The CFO wants outcome data. These are structurally different things, and no amount of better reporting bridges the gap between them.
This post introduces two frameworks that make the problem precise and solvable. The first identifies why AI programs drift toward the wrong metrics. The second gives CIOs and Chief AI Officers a structured way to build evaluation at the level that survives budget scrutiny.
Proxy Metric Trap
The pattern by which enterprise AI programs optimize for measurable proxies of value rather than value itself. A proxy metric is observable and easy to count. Outcome value is real but harder to isolate. Programs fall into the Proxy Metric Trap when the distance between the proxy and the actual outcome grows large enough that improving the proxy no longer improves the business. Queries processed is a proxy. Decisions improved is an outcome. The trap is that the proxy looks like the outcome until a CFO asks.
Outcome Altitude
The level in the business value chain at which an AI program is measured. Altitude 1 is activity (queries processed, documents generated). Altitude 2 is efficiency (time saved, cost per transaction). Altitude 3 is decision quality (better decisions made, errors avoided, risks detected earlier). Altitude 4 is business result (revenue influenced, cost reduction realized, incident rate changed). Most enterprise AI programs operate at Altitude 1 or 2. Boards and CFOs evaluate at Altitude 3 and 4. The gap between where the program measures and where the organization asks is the Evaluation Gap.
The Evaluation Gap is not a data problem. It is a design problem. Programs that close it do so at architecture time, by defining Altitude 3 and 4 success criteria before deployment, not by retrofitting dashboards after the fact.
Why Programs Measure at Altitude 1
The pull toward activity metrics is structural, not a failure of rigor. Activity metrics are available from day one. Every request to an AI system generates a log entry. Query counts, response times, and adoption rates can be reported in the first week of a pilot. They create the appearance of accountability without requiring the organization to agree on what success at Altitude 3 actually looks like.
Altitude 3 and 4 metrics require something harder. They require a baseline. They require agreement on the counterfactual (what would have happened without the AI program). They require tracing influence through a causal chain from AI output to human decision to business result. That causal chain is messy, contested, and takes time to establish. Activity metrics are available immediately and tell a coherent story.
Charles Goodhart's observation from monetary economics applies precisely here [1]. When a measure becomes a target, it ceases to be a good measure. An AI program measured on query volume will be optimized for query volume. Users will route more queries through the system. The system will process them efficiently. The activity metric will improve. The business outcome may not change at all.
The result is an enterprise AI program that performs well on its own scorecard and poorly on the organization's scorecard. That divergence is the Proxy Metric Trap.
The Four-Altitude Architecture
The architecture is not a hierarchy of importance. Activity metrics still matter operationally. A spike in error rates at Altitude 1 is a signal worth acting on. The point is that Altitude 1 and 2 metrics cannot answer the question the organization is actually asking when it decides whether to continue funding an AI program. That question lives at Altitude 3 and 4.
The Evaluation Gap is the structural distance between where the program measures and where the organization asks. A program that reports at Altitude 1 while being evaluated at Altitude 4 has an Evaluation Gap of three levels. That gap does not close on its own. It requires deliberate design.
The Use Case Evaluation Map
Three Failure Modes
Altitude Lock
Programs that begin measuring at Altitude 1 during a pilot rarely climb to Altitude 3 after deployment. The reporting infrastructure was built for activity metrics. The stakeholders who approved the program were shown activity projections. The quarterly business review template asks for adoption and usage. The program is locked at the altitude it launched at, and the Evaluation Gap widens as the organization's expectations grow. Altitude Lock is almost always caused by a measurement architecture decision made in week one that was never revisited.
Counterfactual Blindness
Claiming time savings or error reductions without establishing a pre-deployment baseline. A program that reports "37% faster contract reviews" with no measurement of pre-deployment review times has produced a directional claim, not an outcome measurement. CFOs have seen this pattern in every technology investment cycle. They discount it accordingly. Programs that survive budget scrutiny establish baselines before deployment, not after the question is asked.
Attribution Inflation
Claiming business results that were influenced by many factors, of which the AI program was one. A sales team using an AI-assisted prospecting tool that exceeds its quota by 18% did not produce an 18% lift attributable to the AI program. Market conditions, leadership changes, competitor pricing, and the sales team's own skill all contributed. Programs that claim full attribution for partial influence accelerate CFO skepticism and eventually produce the inverse of what was intended.
Building Evaluation at Altitude 3 and 4
Closing the Evaluation Gap requires three things that most programs skip. They are not technically complex. They are organizationally difficult because they require stakeholder alignment before a line of the program has been written.
The first is a pre-deployment outcome agreement. Before deployment, the program's sponsors and stakeholders agree in writing on what Altitude 3 success looks like for this specific use case. Not generically. Specifically. For a contract review AI, Altitude 3 success might be defined as a measurable reduction in legal escalations from contracts that passed AI review compared to a matched sample from the pre-deployment period. That definition must be agreed before the program launches. Retrofitting it afterward produces contested results.
The second is a baseline measurement period. Four to eight weeks of structured measurement before deployment on the exact metrics the program will eventually claim to improve. Without a baseline, every claim is counterfactual blindness. With a baseline, the program has a credible comparison point.
The third is a causal chain map. A one-page document that traces the path from AI output to human decision to business result. For each step in the chain, it names who makes the decision, what other factors influence it, and how much of the outcome variance is reasonably attributable to AI assistance. This document does not need to be precise. It needs to exist. Its existence forces the organization to be honest about the causal structure of the claim before the claim is made.
The Altitude Assessment for Your Use Cases
Three Enterprise Scenarios
Claims Processing AI with an Altitude 1 Evaluation
A large insurance carrier deploys an AI system to assist claims adjusters with initial claims assessment. At the six-month review, the program reports 12,000 claims assessed with AI assistance, average assessment time reduced from 4.2 to 2.8 hours per claim, and 88% adjuster satisfaction. The CFO asks about claims accuracy, appeals rate, and litigation exposure. The program has no data at those levels. The Evaluation Gap is two altitudes wide.
The business outcome the program was actually meant to support was reduction in claims leakage, the difference between what a carrier pays and what the policy actually requires. Claims leakage is an Altitude 4 metric. The program was deployed without defining a baseline for it. Eighteen months in, the CIO cannot demonstrate that faster assessment improved accuracy or reduced leakage. The program is processing claims faster without evidence of better outcomes.
The fix, had it been applied at design time, was a six-week baseline period measuring claims accuracy, appeals rate, and leakage percentage on the same adjuster population before the AI system was introduced. The program could then have compared those metrics post-deployment and built a credible Altitude 4 case.
Regulatory Monitoring AI with Conflated Attribution
A global bank deploys an AI system to monitor regulatory communications and flag potential compliance obligations. At the annual review, the program reports zero missed regulatory deadlines in the twelve months following deployment, a significant improvement from the three misses in the prior year. The Chief AI Officer presents this as a program success. The CFO notes that the bank also hired two senior compliance analysts and overhauled its regulatory change management process during the same period.
The AI program contributed to improved compliance outcomes. It was not the only contributor. The program has no mechanism to isolate its contribution from the others. The CFO cannot allocate budget to a program whose ROI is entangled with three other investments. Attribution Inflation has made the program's value invisible rather than visible.
The architectural fix is a contribution matrix, agreed at deployment time, that defines which metrics the AI system is the primary driver of versus a contributing factor. The program owns the early-detection latency metric (time from regulatory publication to internal flag). The human analysts own the response quality metric. The process overhaul owns the escalation rate. Each investment has a metric it owns and can demonstrate. The AI system's ROI becomes clear and defensible.
Clinical Documentation AI with No Altitude Agreement
A regional healthcare system deploys an AI system to assist clinicians with documentation. The stated goal at deployment was to reduce clinician burnout and improve documentation quality, enabling better care coordination. Eighteen months in, the program reports significant documentation time savings. A board member asks about care coordination outcomes and readmission rates. The program's evaluation framework was never designed to measure those.
The original goal, reducing burnout and improving care coordination, were Altitude 3 and 4 outcomes. The program measured at Altitude 2. The board's question is legitimate. The documentation time was saved. There is no evidence that the freed time translated to better care coordination or reduced burnout, because neither was measured. The board has a program with significant time savings and no evidence that those savings produced what the deployment was justified to produce.
The pre-deployment outcome agreement discipline would have required the clinical leadership, CIO, and board sponsor to agree in writing on what Altitude 3 success looked like. Agreeing that documentation time savings would be measured AND that clinician burnout survey scores would be tracked quarterly and readmission rates for the relevant patient population would be measured in a comparison cohort. With those baselines in place, the program can answer the board's question.
The Measurement Design Framework
The following four questions must be answered before an enterprise AI program moves from pilot to deployment. They are not a post-deployment checklist. They are deployment gates. A program that cannot answer them is not ready to deploy at scale.
What is the Altitude 3 or 4 outcome this program is meant to influence, stated specifically enough that a CFO could verify it independently?
What is the baseline value of that outcome in the four to eight weeks before this program launches, measured with the same methodology that will be used post-deployment?
What is the causal chain from AI output to the claimed outcome, and what other factors in the same period could produce the same directional change?
Which metrics does this program own exclusively, and which does it influence alongside other investments? The owned metrics are the ones the program will be held to at budget review.
Measurement Architecture by Use Case Type
| Use Case Type | Typical Proxy (Altitude 1-2) | Altitude 3 Metric | Altitude 4 Metric |
|---|---|---|---|
| Document generation | Drafts produced, time per draft | Revision cycles reduced, approval rate | Deal cycle time, negotiation outcome quality |
| Risk and compliance monitoring | Alerts generated, coverage percentage | Relevant alert rate, escalation accuracy | Regulatory incidents avoided, fine exposure reduced |
| Customer service AI | Queries resolved, handle time | First-contact resolution, escalation rate | Customer retention, NPS delta in AI-served cohort |
| Code review and generation | PRs reviewed, time per review | Defect escape rate, security finding rate | Production incident rate, time-to-release |
| Research and analysis | Reports generated, sources synthesized | Decision turnaround time, analysis depth score | Decision quality (measured by outcome of decisions made with vs without AI analysis) |
| Clinical or diagnostic support | Cases reviewed, time per case | Diagnostic accuracy, missed finding rate | Patient outcome metrics in AI-assisted vs matched cohort |
Implementation Roadmap
Outcome Architecture
- Define Altitude 3 and 4 success criteria for each AI use case in the program
- Obtain written alignment from program sponsors, finance, and the relevant business owner
- Identify which metrics the program owns vs influences
- Gate. Alignment document signed before any deployment proceeds
Baseline and Deploy
- Measure Altitude 3 and 4 baselines for six to eight weeks pre-deployment
- Deploy with Altitude 1 and 2 operational monitoring in place
- Begin post-deployment measurement on the same metrics and methodology as the baseline
- Gate. No claims of outcome improvement until four weeks of post-deployment data are available for comparison
Outcome Reporting
- Produce quarterly outcome reports at Altitude 3 and 4 alongside operational dashboards
- Present attribution honestly, naming co-contributors to any claimed business result
- Use the causal chain map to defend contribution claims against CFO and board scrutiny
- Treat each outcome report as a design input for the next program iteration
Cost of Measuring at the Wrong Altitude
Programs that cannot demonstrate Altitude 3 outcomes are the first to be cut in a budget review. Activity metrics do not survive the "so what" question. A program with strong Altitude 1 numbers and no Altitude 3 story is an activity center, not a business investment. Activity centers are discretionary. Business investments are protected.
Enterprise AI programs that cannot demonstrate ROI at Altitude 3 or 4 cannot justify the investment required to scale. The CIO who cannot answer the CFO's outcome question does not get the budget to expand the program from one use case to ten. The Evaluation Gap directly limits the ceiling of the program.
When an AI program consistently reports strong activity metrics without business outcomes, the organization learns to discount AI program reporting entirely. The next program, regardless of actual value, inherits that credibility deficit. Rebuilding institutional trust in AI program measurement takes years and requires demonstrating outcome discipline on multiple deployments in sequence.
Programs measured at Altitude 1 are optimized for Altitude 1. Over time, the program adapts to maximize what it is measured on. A contract review AI measured on volume will process high volumes. It will not necessarily produce better contract outcomes. The program drifts from its original purpose as the optimization target and the business outcome diverge.
Executive Checklist
Yes. It is in writing, signed off by finance and the business owner, and includes a specific measurable outcome and timeframe.
The success criteria were defined during the pilot as whatever the system could demonstrate. They have not been updated for production.
Yes. We measured the relevant Altitude 3 metrics for six to eight weeks before deploying, using the same methodology we use post-deployment.
We estimated the baseline from historical data or vendor benchmarks. We did not measure it directly before deployment.
Yes. We have a causal chain map that names co-contributors. Our program owns specific metrics and influences others alongside different investments.
We present the full business result improvement as the program's outcome. Attribution is not broken down by contributing factor.
Yes. The Altitude 3 and 4 story is in the standard quarterly reporting package and has been rehearsed against CFO-level questions.
We would need to pull additional data. The standard reporting package covers usage and adoption, not business outcomes.
Altitude 3 metrics are included in pilot design. We are building the outcome story during the pilot, not retrofitting it before the deployment review.
Pilots measure feasibility and adoption. We plan to add outcome metrics after we confirm the technology works. This is Altitude Lock in formation.
What Rigorous Evaluation Actually Enables
The purpose of building evaluation at Altitude 3 and 4 is not to make reporting more difficult. It is to make the program defensible at the moments that determine whether it survives and scales. A CFO who understands the causal chain and trusts the attribution methodology is a program's most powerful internal advocate. An AI program that can demonstrate Altitude 4 results with a credible methodology gets the budget to expand. One that cannot gets rationalized in the next cost review.
The BudgetBench evaluation methodology published by Rao and Jaggi [2] demonstrates that rigorous evaluation design applied to AI systems reveals meaningful performance differences that surface only at defined budget tiers and outcome thresholds. The same principle applies at the enterprise program level. Programs evaluated only on activity miss the variation that matters at the level where organizations make investment decisions.
Closing the Evaluation Gap is also the mechanism by which organizations learn what their AI programs actually do. Activity metrics tell you the system is running. Altitude 3 and 4 metrics tell you whether running it is producing what the organization needs. That is the difference between operating a system and understanding one.
For related work on how AI programs maintain vocabulary quality over time and how to design governance for enterprise AI agents, see The Semantic Control Plane and The Action Boundary Gate.
Measure at the altitude where your organization asks questions, not the altitude where your system generates data. Activity data is a diagnostic tool. Outcome data is the investment case. They are not the same thing, and one cannot substitute for the other when budget decisions are being made.
- [1] Goodhart, C.A.E. "Problems of Monetary Management. The U.K. Experience." Papers in Monetary Economics, Reserve Bank of Australia, 1975. The formalization of the principle that once a measure becomes a target, it ceases to function as a reliable measure of the underlying variable.
- [2] Rao, A.K., Jaggi, A. "BudgetBench. A Budget-Tiered Evaluation Protocol for Memory Strategies in Local LLM Agents." arXiv:2609.13149. Demonstrates that evaluation frameworks structured around defined budget tiers and outcome thresholds surface performance differences that activity-level metrics do not detect.