Enterprise AI  ·  Evaluation  ·  ROI

The Evaluation Gap
Why Enterprise AI Programs Measure the Wrong Things

Enterprise AI programs routinely report thousands of queries processed and hours saved on drafts. CFOs and boards ask a different question. Did it change a decision, reduce a risk, or improve a result? Most programs cannot answer that question because they were never designed to.

Arjun Jaggi  ·  September 25, 2026  ·  14 min read
2
Coined Frameworks
4
Outcome Altitudes
3
Enterprise Scenarios

The Budget Conversation That Ends Programs

A CIO presents their AI program results at the quarterly business review. They show query volumes, average response time, user adoption rates, and estimated hours saved based on a self-reported survey. The CFO asks one question. What did it do for the business? The CIO restates the metrics. The CFO schedules a follow-up to discuss budget reallocation.

This is not a communication failure. It is a measurement architecture failure. The program was designed, from the beginning, to measure at the wrong level. The CIO has activity data. The CFO wants outcome data. These are structurally different things, and no amount of better reporting bridges the gap between them.

This post introduces two frameworks that make the problem precise and solvable. The first identifies why AI programs drift toward the wrong metrics. The second gives CIOs and Chief AI Officers a structured way to build evaluation at the level that survives budget scrutiny.

Coined Framework

Proxy Metric Trap

The pattern by which enterprise AI programs optimize for measurable proxies of value rather than value itself. A proxy metric is observable and easy to count. Outcome value is real but harder to isolate. Programs fall into the Proxy Metric Trap when the distance between the proxy and the actual outcome grows large enough that improving the proxy no longer improves the business. Queries processed is a proxy. Decisions improved is an outcome. The trap is that the proxy looks like the outcome until a CFO asks.

Coined Framework

Outcome Altitude

The level in the business value chain at which an AI program is measured. Altitude 1 is activity (queries processed, documents generated). Altitude 2 is efficiency (time saved, cost per transaction). Altitude 3 is decision quality (better decisions made, errors avoided, risks detected earlier). Altitude 4 is business result (revenue influenced, cost reduction realized, incident rate changed). Most enterprise AI programs operate at Altitude 1 or 2. Boards and CFOs evaluate at Altitude 3 and 4. The gap between where the program measures and where the organization asks is the Evaluation Gap.

Core Claim

The Evaluation Gap is not a data problem. It is a design problem. Programs that close it do so at architecture time, by defining Altitude 3 and 4 success criteria before deployment, not by retrofitting dashboards after the fact.

Why Programs Measure at Altitude 1

The pull toward activity metrics is structural, not a failure of rigor. Activity metrics are available from day one. Every request to an AI system generates a log entry. Query counts, response times, and adoption rates can be reported in the first week of a pilot. They create the appearance of accountability without requiring the organization to agree on what success at Altitude 3 actually looks like.

Altitude 3 and 4 metrics require something harder. They require a baseline. They require agreement on the counterfactual (what would have happened without the AI program). They require tracing influence through a causal chain from AI output to human decision to business result. That causal chain is messy, contested, and takes time to establish. Activity metrics are available immediately and tell a coherent story.

Charles Goodhart's observation from monetary economics applies precisely here [1]. When a measure becomes a target, it ceases to be a good measure. An AI program measured on query volume will be optimized for query volume. Users will route more queries through the system. The system will process them efficiently. The activity metric will improve. The business outcome may not change at all.

The result is an enterprise AI program that performs well on its own scorecard and poorly on the organization's scorecard. That divergence is the Proxy Metric Trap.

The Four-Altitude Architecture

ALTITUDE 4 . BUSINESS RESULT Revenue influenced · Cost reduction realized · Incident rate changed · Risk exposure shifted ALTITUDE 3 . DECISION QUALITY Better decisions made · Errors avoided · Risks detected earlier · Judgment improved ALTITUDE 2 . EFFICIENCY Time saved on tasks · Cost per transaction · Throughput increase · Process cycle time ALTITUDE 1 . ACTIVITY Queries processed · Documents generated · Adoption rate · Response time · User count Board asks here CFO asks here Most programs report here Pilots start here and often stay

The architecture is not a hierarchy of importance. Activity metrics still matter operationally. A spike in error rates at Altitude 1 is a signal worth acting on. The point is that Altitude 1 and 2 metrics cannot answer the question the organization is actually asking when it decides whether to continue funding an AI program. That question lives at Altitude 3 and 4.

The Evaluation Gap is the structural distance between where the program measures and where the organization asks. A program that reports at Altitude 1 while being evaluated at Altitude 4 has an Evaluation Gap of three levels. That gap does not close on its own. It requires deliberate design.

The Use Case Evaluation Map

Use Case Evaluation Map  ·  Select a Use Case to See Proxy vs Outcome Metrics

Select any use case above to map its typical proxy metric against the recommended Altitude 3 outcome metric. Adapt this framework to your own AI program evaluation design.

Three Failure Modes

Failure Mode 01

Altitude Lock

Programs that begin measuring at Altitude 1 during a pilot rarely climb to Altitude 3 after deployment. The reporting infrastructure was built for activity metrics. The stakeholders who approved the program were shown activity projections. The quarterly business review template asks for adoption and usage. The program is locked at the altitude it launched at, and the Evaluation Gap widens as the organization's expectations grow. Altitude Lock is almost always caused by a measurement architecture decision made in week one that was never revisited.

Failure Mode 02

Counterfactual Blindness

Claiming time savings or error reductions without establishing a pre-deployment baseline. A program that reports "37% faster contract reviews" with no measurement of pre-deployment review times has produced a directional claim, not an outcome measurement. CFOs have seen this pattern in every technology investment cycle. They discount it accordingly. Programs that survive budget scrutiny establish baselines before deployment, not after the question is asked.

Failure Mode 03

Attribution Inflation

Claiming business results that were influenced by many factors, of which the AI program was one. A sales team using an AI-assisted prospecting tool that exceeds its quota by 18% did not produce an 18% lift attributable to the AI program. Market conditions, leadership changes, competitor pricing, and the sales team's own skill all contributed. Programs that claim full attribution for partial influence accelerate CFO skepticism and eventually produce the inverse of what was intended.

Building Evaluation at Altitude 3 and 4

Closing the Evaluation Gap requires three things that most programs skip. They are not technically complex. They are organizationally difficult because they require stakeholder alignment before a line of the program has been written.

The first is a pre-deployment outcome agreement. Before deployment, the program's sponsors and stakeholders agree in writing on what Altitude 3 success looks like for this specific use case. Not generically. Specifically. For a contract review AI, Altitude 3 success might be defined as a measurable reduction in legal escalations from contracts that passed AI review compared to a matched sample from the pre-deployment period. That definition must be agreed before the program launches. Retrofitting it afterward produces contested results.

The second is a baseline measurement period. Four to eight weeks of structured measurement before deployment on the exact metrics the program will eventually claim to improve. Without a baseline, every claim is counterfactual blindness. With a baseline, the program has a credible comparison point.

The third is a causal chain map. A one-page document that traces the path from AI output to human decision to business result. For each step in the chain, it names who makes the decision, what other factors influence it, and how much of the outcome variance is reasonably attributable to AI assistance. This document does not need to be precise. It needs to exist. Its existence forces the organization to be honest about the causal structure of the claim before the claim is made.

The Altitude Assessment for Your Use Cases

Evaluation Altitude by AI Use Case Type
Directional illustration. Bars show typical reported altitude vs recommended measurement altitude. Based on practitioner observation across enterprise AI deployments.

Three Enterprise Scenarios

Scenario 01  ·  Insurance Carrier  ·  CIO

Claims Processing AI with an Altitude 1 Evaluation

A large insurance carrier deploys an AI system to assist claims adjusters with initial claims assessment. At the six-month review, the program reports 12,000 claims assessed with AI assistance, average assessment time reduced from 4.2 to 2.8 hours per claim, and 88% adjuster satisfaction. The CFO asks about claims accuracy, appeals rate, and litigation exposure. The program has no data at those levels. The Evaluation Gap is two altitudes wide.

The business outcome the program was actually meant to support was reduction in claims leakage, the difference between what a carrier pays and what the policy actually requires. Claims leakage is an Altitude 4 metric. The program was deployed without defining a baseline for it. Eighteen months in, the CIO cannot demonstrate that faster assessment improved accuracy or reduced leakage. The program is processing claims faster without evidence of better outcomes.

The fix, had it been applied at design time, was a six-week baseline period measuring claims accuracy, appeals rate, and leakage percentage on the same adjuster population before the AI system was introduced. The program could then have compared those metrics post-deployment and built a credible Altitude 4 case.

Scenario 02  ·  Global Bank  ·  Chief AI Officer

Regulatory Monitoring AI with Conflated Attribution

A global bank deploys an AI system to monitor regulatory communications and flag potential compliance obligations. At the annual review, the program reports zero missed regulatory deadlines in the twelve months following deployment, a significant improvement from the three misses in the prior year. The Chief AI Officer presents this as a program success. The CFO notes that the bank also hired two senior compliance analysts and overhauled its regulatory change management process during the same period.

The AI program contributed to improved compliance outcomes. It was not the only contributor. The program has no mechanism to isolate its contribution from the others. The CFO cannot allocate budget to a program whose ROI is entangled with three other investments. Attribution Inflation has made the program's value invisible rather than visible.

The architectural fix is a contribution matrix, agreed at deployment time, that defines which metrics the AI system is the primary driver of versus a contributing factor. The program owns the early-detection latency metric (time from regulatory publication to internal flag). The human analysts own the response quality metric. The process overhaul owns the escalation rate. Each investment has a metric it owns and can demonstrate. The AI system's ROI becomes clear and defensible.

Scenario 03  ·  Healthcare System  ·  Board

Clinical Documentation AI with No Altitude Agreement

A regional healthcare system deploys an AI system to assist clinicians with documentation. The stated goal at deployment was to reduce clinician burnout and improve documentation quality, enabling better care coordination. Eighteen months in, the program reports significant documentation time savings. A board member asks about care coordination outcomes and readmission rates. The program's evaluation framework was never designed to measure those.

The original goal, reducing burnout and improving care coordination, were Altitude 3 and 4 outcomes. The program measured at Altitude 2. The board's question is legitimate. The documentation time was saved. There is no evidence that the freed time translated to better care coordination or reduced burnout, because neither was measured. The board has a program with significant time savings and no evidence that those savings produced what the deployment was justified to produce.

The pre-deployment outcome agreement discipline would have required the clinical leadership, CIO, and board sponsor to agree in writing on what Altitude 3 success looked like. Agreeing that documentation time savings would be measured AND that clinician burnout survey scores would be tracked quarterly and readmission rates for the relevant patient population would be measured in a comparison cohort. With those baselines in place, the program can answer the board's question.

The Measurement Design Framework

The following four questions must be answered before an enterprise AI program moves from pilot to deployment. They are not a post-deployment checklist. They are deployment gates. A program that cannot answer them is not ready to deploy at scale.

Question 01

What is the Altitude 3 or 4 outcome this program is meant to influence, stated specifically enough that a CFO could verify it independently?

Question 02

What is the baseline value of that outcome in the four to eight weeks before this program launches, measured with the same methodology that will be used post-deployment?

Question 03

What is the causal chain from AI output to the claimed outcome, and what other factors in the same period could produce the same directional change?

Question 04

Which metrics does this program own exclusively, and which does it influence alongside other investments? The owned metrics are the ones the program will be held to at budget review.

Measurement Architecture by Use Case Type

Use Case Type Typical Proxy (Altitude 1-2) Altitude 3 Metric Altitude 4 Metric
Document generation Drafts produced, time per draft Revision cycles reduced, approval rate Deal cycle time, negotiation outcome quality
Risk and compliance monitoring Alerts generated, coverage percentage Relevant alert rate, escalation accuracy Regulatory incidents avoided, fine exposure reduced
Customer service AI Queries resolved, handle time First-contact resolution, escalation rate Customer retention, NPS delta in AI-served cohort
Code review and generation PRs reviewed, time per review Defect escape rate, security finding rate Production incident rate, time-to-release
Research and analysis Reports generated, sources synthesized Decision turnaround time, analysis depth score Decision quality (measured by outcome of decisions made with vs without AI analysis)
Clinical or diagnostic support Cases reviewed, time per case Diagnostic accuracy, missed finding rate Patient outcome metrics in AI-assisted vs matched cohort

Implementation Roadmap

Phase 1  ·  Weeks 1 to 4

Outcome Architecture

  • Define Altitude 3 and 4 success criteria for each AI use case in the program
  • Obtain written alignment from program sponsors, finance, and the relevant business owner
  • Identify which metrics the program owns vs influences
  • Gate. Alignment document signed before any deployment proceeds
Phase 2  ·  Weeks 5 to 12

Baseline and Deploy

  • Measure Altitude 3 and 4 baselines for six to eight weeks pre-deployment
  • Deploy with Altitude 1 and 2 operational monitoring in place
  • Begin post-deployment measurement on the same metrics and methodology as the baseline
  • Gate. No claims of outcome improvement until four weeks of post-deployment data are available for comparison
Phase 3  ·  Week 13 Forward

Outcome Reporting

  • Produce quarterly outcome reports at Altitude 3 and 4 alongside operational dashboards
  • Present attribution honestly, naming co-contributors to any claimed business result
  • Use the causal chain map to defend contribution claims against CFO and board scrutiny
  • Treat each outcome report as a design input for the next program iteration

Cost of Measuring at the Wrong Altitude

Budget Vulnerability

Programs that cannot demonstrate Altitude 3 outcomes are the first to be cut in a budget review. Activity metrics do not survive the "so what" question. A program with strong Altitude 1 numbers and no Altitude 3 story is an activity center, not a business investment. Activity centers are discretionary. Business investments are protected.

Scaling Blocked

Enterprise AI programs that cannot demonstrate ROI at Altitude 3 or 4 cannot justify the investment required to scale. The CIO who cannot answer the CFO's outcome question does not get the budget to expand the program from one use case to ten. The Evaluation Gap directly limits the ceiling of the program.

Trust Erosion

When an AI program consistently reports strong activity metrics without business outcomes, the organization learns to discount AI program reporting entirely. The next program, regardless of actual value, inherits that credibility deficit. Rebuilding institutional trust in AI program measurement takes years and requires demonstrating outcome discipline on multiple deployments in sequence.

Strategic Misalignment

Programs measured at Altitude 1 are optimized for Altitude 1. Over time, the program adapts to maximize what it is measured on. A contract review AI measured on volume will process high volumes. It will not necessarily produce better contract outcomes. The program drifts from its original purpose as the optimization target and the business outcome diverge.

Executive Checklist

Does your AI program have a documented Altitude 3 or 4 success definition, agreed before deployment?
Good answer

Yes. It is in writing, signed off by finance and the business owner, and includes a specific measurable outcome and timeframe.

Red flag

The success criteria were defined during the pilot as whatever the system could demonstrate. They have not been updated for production.

Does your program have a pre-deployment baseline for the outcomes it claims to improve?
Good answer

Yes. We measured the relevant Altitude 3 metrics for six to eight weeks before deploying, using the same methodology we use post-deployment.

Red flag

We estimated the baseline from historical data or vendor benchmarks. We did not measure it directly before deployment.

Can you explain what other factors contributed to any business result your program claims to have influenced?
Good answer

Yes. We have a causal chain map that names co-contributors. Our program owns specific metrics and influences others alongside different investments.

Red flag

We present the full business result improvement as the program's outcome. Attribution is not broken down by contributing factor.

Could your CIO answer a CFO's outcome question in a 10-minute budget meeting without additional preparation?
Good answer

Yes. The Altitude 3 and 4 story is in the standard quarterly reporting package and has been rehearsed against CFO-level questions.

Red flag

We would need to pull additional data. The standard reporting package covers usage and adoption, not business outcomes.

Does your pilot evaluation design include Altitude 3 metrics, or only Altitude 1 and 2?
Good answer

Altitude 3 metrics are included in pilot design. We are building the outcome story during the pilot, not retrofitting it before the deployment review.

Red flag

Pilots measure feasibility and adoption. We plan to add outcome metrics after we confirm the technology works. This is Altitude Lock in formation.

What Rigorous Evaluation Actually Enables

The purpose of building evaluation at Altitude 3 and 4 is not to make reporting more difficult. It is to make the program defensible at the moments that determine whether it survives and scales. A CFO who understands the causal chain and trusts the attribution methodology is a program's most powerful internal advocate. An AI program that can demonstrate Altitude 4 results with a credible methodology gets the budget to expand. One that cannot gets rationalized in the next cost review.

The BudgetBench evaluation methodology published by Rao and Jaggi [2] demonstrates that rigorous evaluation design applied to AI systems reveals meaningful performance differences that surface only at defined budget tiers and outcome thresholds. The same principle applies at the enterprise program level. Programs evaluated only on activity miss the variation that matters at the level where organizations make investment decisions.

Closing the Evaluation Gap is also the mechanism by which organizations learn what their AI programs actually do. Activity metrics tell you the system is running. Altitude 3 and 4 metrics tell you whether running it is producing what the organization needs. That is the difference between operating a system and understanding one.

For related work on how AI programs maintain vocabulary quality over time and how to design governance for enterprise AI agents, see The Semantic Control Plane and The Action Boundary Gate.

Key Takeaway

Measure at the altitude where your organization asks questions, not the altitude where your system generates data. Activity data is a diagnostic tool. Outcome data is the investment case. They are not the same thing, and one cannot substitute for the other when budget decisions are being made.

References

Excited about AI, innovation, and growth?

Start a conversation