↓ Download PDF
Arjun Jaggi
Enterprise AI Research · White Paper No. 02
A FLUENT ANSWER UNROOTED, UNVERIFIABLE THE DEFENSIBILITY GAP ROOTED IN EVIDENCE SOURCED · REVIEWED · REPRODUCIBLE
October 2026
The Defensibility
Gap
Local open-weight models, and the evidence layer that makes their answers defensible
Authors
Arjun Jaggi
Aditya Karnam Gururaj Rao
Framework
DDF, Version 1.0
Discipline
Defensible AI Discovery
Readership
Boards · C-suite · Research leadership
The Defensibility GapArjun Jaggi & Aditya Karnam Gururaj Rao
Contents
·
Foreword
A letter to the reader
03
·
Executive Summary
Five findings for the board
04
01
The Problem and the Market Need
Why a fluent answer is not a defensible one
05
02
The Anatomy of an Indefensible Answer
Six failure modes an answer must survive
06
03
The Evidence Layer
Five components, one defensible answer
07
04
The Economics
The inversion and the cost of being wrong
08
05
The Implementation Roadmap
Fifteen weeks, two hard gates
10
06
The Decision Framework
Four variables, four deployment postures
11
·
The DDF on a Page
The complete framework, one spread, desk-ready
12
·
Self-Assessment Scorecard
Ten questions place your program on the maturity curve
13
·
Methodology, Notes, and the Author
14

About This Paper

This paper introduces the Defensible Discovery Framework (DDF), a method for turning the fluent answers of a local open-weight model into answers an organization can defend. It is written for the leaders now moving real research, analysis, and consequential decisions onto models they run themselves, and who need those answers to survive an auditor, a board, or a peer.

The framework, the ten exhibits, and the terms Defensibility Gap and Evidence Debt are original to this work, published under the license on the back cover. The evidence primitives it builds on, pre-registration and the Kill Threshold among them, are drawn from the research agent Zorp and cited where used. Every quantitative claim is cited to a named source or labeled directional, without exception.

How to read the exhibits
Solid fill, cited data
Hatched fill, directional illustration
Red, the defensibility gap
Arjun Jaggi · Aditya Karnam Gururaj Rao02
The Defensibility GapForeword
Foreword
A letter to the reader

The first time a model answered a question we could not have answered ourselves, we believed it. The second time, it was wrong, fluently and completely, and from the outside the two occasions were identical. That is the problem this paper is about, stated in one experience. The machine had become good enough to be trusted and not yet accountable enough to deserve it.

For a decade the hard part was the model. That era is ending. On a public leaderboard the gap between the best open-weight model and the best closed one fell from roughly eight percent to under two in a single year [1]. A capable model now runs on a laptop. When everyone can get a fluent answer, the fluent answer stops being worth anything. What becomes scarce, and valuable, is an answer you can defend.

This paper is our account of what that evidence looks like and how to build it. Not a better model, but a layer around the model that records what it assumed, what it read, what it found that pointed the other way, and what would have proved it wrong. The framework is original work. The numbers are cited or labeled directional, without exception. We hold this to the standard of the rooms it will be read in.

Arjun Jaggi · Aditya Karnam Gururaj Rao
Enterprise AI Research · arjunjaggi.com

At a Glance

1.70%
Capability gap, best open-weight vs best closed model, Feb 2025, down from 8.04% a year earlier [1]
98%
Best-case inference cost cut while matching a frontier model's quality [3]
6
Failure modes a defensible answer must close
15 wks
From fluent answers to a governed defensible capability, directional
Arjun Jaggi · Aditya Karnam Gururaj Rao03
The Defensibility GapExecutive Summary
Executive Summary
Five findings for the board

The cost of a capable model has collapsed, and with it the advantage of having one. A local open-weight model will now answer almost any question an organization can ask. What it will not do, on its own, is make that answer defensible. Five findings follow.

1

The model stopped being the moat. The gap between the best open-weight and the best closed model fell from 8.04 percent to 1.70 percent in thirteen months [1]. When a capable model runs on commodity hardware, capability is no longer the differentiator. What an organization can defend is.

2

A fluent answer is not a defensible one. A model produces a confident answer to a hard question in seconds. It does not, unprompted, state what it assumed, what evidence it weighed, or what it found that pointed the other way. The Defensibility Gap is the distance between those two answers.

3

Six failure modes make an answer indefensible. The DDF taxonomy names each one and maps it to the evidence control that closes it. A gap left in any single mode is the opening a challenge will find first, which is why partial evidence fails.

4

The evidence layer is buildable, local, and measurable. It is not a model and not a data center. It is a set of controls that run on open-weight models on commodity hardware, and whose rigor can be measured under a fixed budget rather than asserted [2].

5

Defensible discovery is a one-quarter build. A team of four to five takes a single workflow from fluent answers to a governed, defensible capability in roughly fifteen weeks, directionally, with no dependency on a frontier vendor.

Arjun Jaggi · Aditya Karnam Gururaj Rao04
The Defensibility GapSection 01
01
Section 01 · The Problem
The hard part moved from the model to the evidence

For a decade the binding constraint on applied AI was access to a capable model. That constraint is gone. On a public leaderboard the lead of the best closed model over the best open-weight model fell from 8.04 percent to 1.70 percent between January 2024 and February 2025 [1], and an open-weight model now runs on hardware an analyst already owns. The scarce thing is no longer the answer. It is the ability to defend it.

A model operates on fluency. A defensible answer operates on evidence. A fluent answer arrives complete, confident, and without a trace of how it was reached, and nothing on its surface tells you whether it is right. The research agent Zorp, built for this exact problem, states it plainly. A confident answer is not a defensible one [5].

Framework Term

Defensibility Gap. The distance between an answer a model will produce in seconds and one that survives a challenge. Every component of this framework exists to close it.

EXHIBIT 1
The model gap nearly closed in a year, so the model stopped being the edge
Lead of best closed over best open-weight model, Chatbot Arena, in points
0 5% 10% 8.04% 1.70% JAN 2024 FEB 2025 -6.34 PTS
Source. Stanford HAI, Artificial Intelligence Index Report, 2025 [1].

The structural point is shown in Exhibit 2. A question does not become a defensible answer by being answered. It becomes one only by passing through an evidence layer that most deployments skip entirely, which is where the gap opens.

Arjun Jaggi · Aditya Karnam Gururaj Rao05
The Defensibility GapSection 02
02
Section 02 · The Taxonomy
Six ways a confident answer fails to be defensible

An answer is indefensible in specific, nameable ways, not vaguely. Each failure mode below is a question a challenger will ask, and each maps to one control in the evidence layer of Section 03. A gap left in any single mode is the opening a challenge finds first.

I
Unstated Assumptions

The answer rests on premises it never names, so a reader cannot test the one premise that actually carried the conclusion.

II
Selective Evidence

Only the sources that agree are cited. The search stopped when the answer appeared, not when the evidence was exhausted.

III
No Falsification Criterion

Nothing was set in advance that would have proved the answer wrong, so the investigation could not have failed and the result means less than it looks.

IV
Buried Conflict

Evidence that pointed the other way exists and is not shown. The reader is handed a verdict, not the disagreement behind it.

V
Unreproducible Path

The reasoning cannot be re-run. A second attempt, or an auditor, cannot reach the same place by the same route.

VI
Absent Provenance

Nothing records what the answer depends on, so when a source is later retracted, no one can trace which conclusions fall with it.

EXHIBIT 2
The defensible answer is produced after the gap, not before it
FLUENT ANSWER Produced in seconds, undefended WHERE THE FAILURE MODES LIVE I Assumptions II Selective evidence III No falsification IV Buried conflict V Unreproducible VI No provenance THE DEFENSIBILITY GAP THE EVIDENCE LAYER Five components, Section 03 DEFENSIBLE ANSWER EACH COMPONENT CLOSES ONE FAILURE MODE Same question, same model. The only difference between the two answers is the layer on the right.
Source. DDF taxonomy, original to this work. The six modes are exhaustive for single-answer discovery tasks; multi-agent settings add orchestration failure, out of scope here.
Framework Term

Evidence Debt. The unpaid cost of shipping answers you cannot defend. It accrues silently with every undocumented decision and comes due in full the first time one is challenged.

Arjun Jaggi · Aditya Karnam Gururaj Rao06
The Defensibility GapSection 03
03
Section 03 · The Architecture
Five components turn a fluent answer into a defensible one

The evidence layer is a pipeline, not a product. A question enters, and a defensible answer and a durable artifact leave. Each of the five components produces one output a challenger can inspect, and each closes one of the six failure modes from Section 02, all on an open-weight model the team runs itself.

EXHIBIT 3
The evidence layer is five inspectable steps between a question and an answer
QUESTION C1 PRE-REGISTER claim + threshold registered claim C2 INVESTIGATE sources to a rule source ledger C3 SURFACE CONFLICT conflict log C4 ADVERSARIAL REVIEW review verdict C5 PROVENANCE reproducible evidence artifact DEFENSIBLE ANSWER THE EVIDENCE LAYER · RUNS LOCALLY ON OPEN-WEIGHT MODELS
Source. DDF evidence layer, original to this work. The pre-registration and adversarial-review primitives are instantiated in the research agent Zorp [5], which runs the full layer on a local model.
Failure modeClosed byMechanismInspectable output
I. Unstated AssumptionsC1 Pre-RegisterClaim and assumptions written and hashed before any evidence is gatheredRegistered claim
II. Selective EvidenceC2 InvestigateSearch runs to a stopping rule, every source recorded whether it agrees or notSource ledger
III. No FalsificationC1 Pre-RegisterA Kill Threshold set in advance states what would prove the answer wrongFalsification criterion
IV. Buried ConflictC3 Surface ConflictDisconfirming evidence is reported in the answer, not discarded on the wayConflict log
V. Unreproducible PathC5 ProvenanceThe full reasoning trace is captured so a second run reaches the same placeReproducible trace
VI. Absent ProvenanceC5 ProvenanceEvery conclusion is linked to the evidence it depends on, retraction-awareEvidence artifact
Arjun Jaggi · Aditya Karnam Gururaj Rao07
The Defensibility GapSection 04
04
Section 04 · The Economics
The cost moved from the model to the mistake

The economics of defensible discovery invert the ones most organizations priced for. A hosted frontier model charges for every question, forever. A local open-weight model is paid for once, in hardware and setup, and answers the next question at a marginal cost that approaches zero. Published work shows a frontier model's quality can be matched at up to 98 percent lower cost in the best case [3]. Past a modest volume, local is simply cheaper.

But the inference bill is not the cost that matters. The cost that matters is being confidently wrong. The Defensibility Gap is paid not in cents per query but in the single decision a fluent, unexamined answer sends in the wrong direction, and that bill arrives whether the answer was hosted or local. Evidence is the only thing that lowers it.

98%
The best-case reduction in inference cost while matching a frontier model's quality, once the work runs on a local model [3]
EXHIBIT 4
Hosted cost scales with every question, so past break-even local is cheaper
Directional illustration. Cumulative cost as questions answered rises
$0 high LOCAL CHEAPER BREAK-EVEN Build cost paid once HOSTED FRONTIER API LOCAL + EVIDENCE LAYER QUESTIONS ANSWERED OVER TIME
Source. Directional illustration, original to this work. The hosted curve rises with usage; the local curve is dominated by a one-time build. Cost-parity evidence for local operation draws on FrugalGPT [3].
Arjun Jaggi · Aditya Karnam Gururaj Rao08
The Defensibility GapSection 04 · Continued
EXHIBIT 5
No single component closes every failure mode, which is the case for the full layer
Coverage of each failure mode by each component, directional, in percent
C1 PRE-REG C2 INVEST C3 CONFLICT C4 REVIEW C5 PROV I. Unstated Assumptions II. Selective Evidence III. No Falsification IV. Buried Conflict V. Unreproducible Path VI. Absent Provenance 90 20 10 55 25 15 90 40 55 20 95 10 15 60 20 10 35 90 60 25 20 30 15 45 90 15 40 20 45 95 LOW HIGH C4 review is the backstop across every row
Source. Directional illustration from pilot observation, original to this work. Column C4 carries moderate coverage in every row, so review is a backstop and never a substitute for the primary control.
EXHIBIT 6
A wrong decision costs more than the wrong answer
Relative composition of the cost of an indefensible decision, directional
WRONG ANSWER REWORK DELAY CREDIBILITY TOTAL ALL SEGMENTS DIRECTIONAL, NOT TO SCALE
Source. Directional illustration, original to this work. No single dollar figure anchors the cost of a wrong decision, which varies by the decision; the shape, not the scale, is the point.

The Return on Control

Cost of inaction. A wrong decision made on an unexamined answer carries rework, delay, and the loss of credibility when it surfaces. None of it lands as an AI line item. All of it is Evidence Debt coming due.

Cost of control. The evidence layer is a six-to-twelve-week build for four to five people, on hardware already owned. It runs on the same local model the answers already use, so it adds rigor, not infrastructure.

Payback logic. One avoided wrong decision of consequence exceeds the build cost. The layer does not need to catch many. It needs to catch the one that would otherwise have shipped.

Arjun Jaggi · Aditya Karnam Gururaj Rao09
The Defensibility GapSection 05
05
Section 05 · The Roadmap
Fifteen weeks, three phases, two hard gates

Each gate is a measurable exit condition, not a status meeting. A program that cannot make one workflow defensible in six weeks has no business extending the method across an estate. The pilot is deliberately small, because the discipline, not the scale, is what has to be proven first.

EXHIBIT 7
The timeline front-loads pre-registration, the control Exhibit 5 shows carries the most
WK 1-2 WK 3-4 WK 5-6 WK 7-10 WK 11-14 WK 15+ PHASE 1 Pilot PHASE 2 Harden PHASE 3 Scale Pick workflow + stand up local model Write Kill Thresholds Source ledger + pilot run GATE 1 Conflict surfacing + review panel Adversarial review Provenance artifacts GATE 2 All workflows, estate-wide Reproducibility + quarterly audit Phase 1 Phase 2 Phase 3 Go / no-go gate
Source. DDF deployment model, original to this work. Pre-registration lands in weeks 1 to 4 because it is the control with the widest coverage in Exhibit 5.
1
Weeks 1-6
Pilot and pre-register
  • Select one high-stakes workflow
  • Stand up a local open-weight model
  • Write Kill Thresholds for the first claims
  • Build the source ledger
  • Gate 1. Every pilot answer carries a registered claim and a threshold
2
Weeks 7-14
Harden the layer
  • Add conflict surfacing
  • Stand up an adversarial review panel
  • Emit a provenance artifact per answer
  • Check each answer against its threshold
  • Gate 2. Zero answers released without an evidence artifact
3
Weeks 15+
Scale defensibly
  • Extend to every consequential workflow
  • Automate reproducibility checks
  • Set a quarterly adversarial audit
  • Publish the internal evidence standard
  • Success. A challenged answer reconstructs from its artifact alone
Arjun Jaggi · Aditya Karnam Gururaj Rao10
The Defensibility GapSection 06
06
Section 06 · The Decision
Four variables select the posture, and the posture sets the layer
1
Decision Stakes

A reversible draft tolerates a fluent answer. An irreversible or externally visible decision demands the full layer, regardless of how good the model is.

2
Scrutiny

Who will challenge this, an auditor, a regulator, a board, a peer? The more adversarial the audience, the more of the evidence layer the answer has to carry.

3
Data Sensitivity

Can the question and its sources leave your environment? When they cannot, the model must be local, which the closing capability gap now makes practical.

4
Reuse Horizon

Acted on once, or relied on for months? A durable conclusion needs provenance so that a later retraction can be traced to the decisions it touched.

EXHIBIT 8
Three questions resolve any discovery task to one of four postures
DISCOVERY TASK HIGH STAKES? CHALLENGED? REPRODUCED? YES YES NO NO NO YES POSTURE A Draft assist no evidence layer POSTURE B Private analysis local, light evidence POSTURE C Reviewed decision local, full review POSTURE D Published finding full layer + artifact
Source. DDF posture model, original to this work. An organization runs mixed postures by design; a Posture A brainstorm and a Posture D published finding can share the same local model.
EXHIBIT 9
What each posture carries, at a glance
POSTURE C1 C2 C3 C4 C5 ADDED CONTROL A · Draft assist None required B · Private analysis Pre-register + source C · Reviewed decision Adversarial review D · Published finding Reproducible artifact
Source. DDF posture model, original to this work. Filled circles mark required components; the posture sets the floor, never the ceiling.

The closing argument. A capable model is now a commodity, and a commodity confers no advantage. The organizations that pull ahead will be the ones whose answers can be trusted, challenged, and reproduced, on models they run themselves. That trust is what the evidence layer builds, and the fifteen weeks start whenever you do.

Arjun Jaggi · Aditya Karnam Gururaj Rao11
The Defensibility GapThe DDF on a Page
The DDF on a Page
Pin This Page
01 · Six failure modes of an undefended answer
I
Unstated Assumptions
II
Selective Evidence
III
No Falsification
IV
Buried Conflict
V
Unreproducible Path
VI
Absent Provenance
02 · The evidence layer, five inspectable steps
C1
Pre-register
›
C2
Investigate
›
C3
Surface conflict
›
C4
Review
›
C5
Provenance
03 · Four deployment postures
A Draft assist
No layer. Brainstorms and drafts.
B Private analysis
Local, pre-register and source.
C Reviewed decision
Local, through adversarial review.
D Published finding
Full layer plus reproducible artifact.
04 · Fifteen weeks, two hard gates
Phase 1 · Pilot and pre-register
GATE 1
Phase 2 · Harden the layer
GATE 2
Phase 3 · Scale
Framework vocabulary
Defensibility Gap Evidence Debt Kill Threshold · Zorp The Evidence Layer
Answers are cheap. Evidence is not.
Pin this page, then run the
scorecard on page 13 to place your program
Arjun Jaggi · Aditya Karnam Gururaj Rao12
The Defensibility GapSelf-Assessment
Self-Assessment · Ten Questions
Where does your program stand?

Answer honestly for your most consequential AI-assisted workflow. Count the YES answers, then read your level on the band below and the component gaps it points to. Designed to be completed in pen.

No. Tests Question Yes No
1C1Before a consequential answer, is the claim and its assumptions written down in advance?
2C1Is a falsification criterion, what would prove the answer wrong, set before the work starts?
3C2Does the investigation record every source it consulted, including the ones that disagreed?
4C2Does the search run to a stopping rule rather than stopping the moment an answer appears?
5C3Is disconfirming evidence reported in the final answer, not discarded along the way?
6C4Does an adversarial reviewer challenge the answer before it is released?
7C4Is every answer checked against its own falsification criterion before release?
8C5Can a second person reproduce the answer from the recorded reasoning alone?
9C5Is every conclusion linked to the evidence it depends on, so a retraction can be traced?
10LOCALDoes consequential work run on a model you control, so sensitive questions never leave?
Your score, your maturity level
0-2 YES
Level 1 · Fluent
3-4 YES
Level 2 · Sourced
5-6 YES
Level 3 · Reviewed
7-8 YES
Level 4 · Reproducible
9-10 YES
Level 5 · Defensible
Arjun Jaggi · Aditya Karnam Gururaj Rao13
The Defensibility GapMethodology · Notes · Author

About the Research

This paper is a conceptual contribution. The DDF taxonomy, the five-component evidence layer, the four-posture model, and the terms Defensibility Gap and Evidence Debt are original to this work. The pre-registration and adversarial-review primitives it builds on are instantiated in the research agent Zorp and cited where used. Holistic evaluation of language models established that quality is multidimensional and not captured by any single score [4]; the evidence layer extends that principle from measuring a model to defending an answer.

Quantitative claims follow one discipline throughout, made visible in every exhibit. Solid fills carry cited data. Hatched fills and dashed curves labeled directional exist to make the shape of an argument legible, not to report measurements. No statistic in this paper is attributed to a source that cannot be independently verified.

Notes

About the Authors

AJ
Arjun Jaggi and Aditya Karnam Gururaj Rao

Arjun Jaggi and Aditya Karnam Gururaj Rao write and advise at the intersection of enterprise AI strategy, governance, and defensible discovery. Their published body of work, spanning original frameworks, concept papers, and books, is read by executive teams navigating AI adoption and is available in full, without a paywall, at arjunjaggi.com.

Engage · arjunjaggi.com · calendly.com/arjunjaggi
Arjun Jaggi · Aditya Karnam Gururaj Rao14
Arjun Jaggi
Enterprise AI Research
Answers are cheap. Evidence is not. The gap between the two is where the next decade of enterprise AI will be won.
Defensibility Gap Evidence Debt The Evidence Layer DDF
© 2026 Arjun Jaggi and Aditya Karnam Gururaj Rao. All rights reserved. Academic citation permitted with attribution; commercial use and derivative frameworks require written permission.
White Paper No. 02