This paper introduces the Defensible Discovery Framework (DDF), a method for turning the fluent answers of a local open-weight model into answers an organization can defend. It is written for the leaders now moving real research, analysis, and consequential decisions onto models they run themselves, and who need those answers to survive an auditor, a board, or a peer.
The framework, the ten exhibits, and the terms Defensibility Gap and Evidence Debt are original to this work, published under the license on the back cover. The evidence primitives it builds on, pre-registration and the Kill Threshold among them, are drawn from the research agent Zorp and cited where used. Every quantitative claim is cited to a named source or labeled directional, without exception.
The first time a model answered a question we could not have answered ourselves, we believed it. The second time, it was wrong, fluently and completely, and from the outside the two occasions were identical. That is the problem this paper is about, stated in one experience. The machine had become good enough to be trusted and not yet accountable enough to deserve it.
For a decade the hard part was the model. That era is ending. On a public leaderboard the gap between the best open-weight model and the best closed one fell from roughly eight percent to under two in a single year [1]. A capable model now runs on a laptop. When everyone can get a fluent answer, the fluent answer stops being worth anything. What becomes scarce, and valuable, is an answer you can defend.
This paper is our account of what that evidence looks like and how to build it. Not a better model, but a layer around the model that records what it assumed, what it read, what it found that pointed the other way, and what would have proved it wrong. The framework is original work. The numbers are cited or labeled directional, without exception. We hold this to the standard of the rooms it will be read in.
The cost of a capable model has collapsed, and with it the advantage of having one. A local open-weight model will now answer almost any question an organization can ask. What it will not do, on its own, is make that answer defensible. Five findings follow.
The model stopped being the moat. The gap between the best open-weight and the best closed model fell from 8.04 percent to 1.70 percent in thirteen months [1]. When a capable model runs on commodity hardware, capability is no longer the differentiator. What an organization can defend is.
A fluent answer is not a defensible one. A model produces a confident answer to a hard question in seconds. It does not, unprompted, state what it assumed, what evidence it weighed, or what it found that pointed the other way. The Defensibility Gap is the distance between those two answers.
Six failure modes make an answer indefensible. The DDF taxonomy names each one and maps it to the evidence control that closes it. A gap left in any single mode is the opening a challenge will find first, which is why partial evidence fails.
The evidence layer is buildable, local, and measurable. It is not a model and not a data center. It is a set of controls that run on open-weight models on commodity hardware, and whose rigor can be measured under a fixed budget rather than asserted [2].
Defensible discovery is a one-quarter build. A team of four to five takes a single workflow from fluent answers to a governed, defensible capability in roughly fifteen weeks, directionally, with no dependency on a frontier vendor.
For a decade the binding constraint on applied AI was access to a capable model. That constraint is gone. On a public leaderboard the lead of the best closed model over the best open-weight model fell from 8.04 percent to 1.70 percent between January 2024 and February 2025 [1], and an open-weight model now runs on hardware an analyst already owns. The scarce thing is no longer the answer. It is the ability to defend it.
A model operates on fluency. A defensible answer operates on evidence. A fluent answer arrives complete, confident, and without a trace of how it was reached, and nothing on its surface tells you whether it is right. The research agent Zorp, built for this exact problem, states it plainly. A confident answer is not a defensible one [5].
Defensibility Gap. The distance between an answer a model will produce in seconds and one that survives a challenge. Every component of this framework exists to close it.
The structural point is shown in Exhibit 2. A question does not become a defensible answer by being answered. It becomes one only by passing through an evidence layer that most deployments skip entirely, which is where the gap opens.
An answer is indefensible in specific, nameable ways, not vaguely. Each failure mode below is a question a challenger will ask, and each maps to one control in the evidence layer of Section 03. A gap left in any single mode is the opening a challenge finds first.
The answer rests on premises it never names, so a reader cannot test the one premise that actually carried the conclusion.
Only the sources that agree are cited. The search stopped when the answer appeared, not when the evidence was exhausted.
Nothing was set in advance that would have proved the answer wrong, so the investigation could not have failed and the result means less than it looks.
Evidence that pointed the other way exists and is not shown. The reader is handed a verdict, not the disagreement behind it.
The reasoning cannot be re-run. A second attempt, or an auditor, cannot reach the same place by the same route.
Nothing records what the answer depends on, so when a source is later retracted, no one can trace which conclusions fall with it.
Evidence Debt. The unpaid cost of shipping answers you cannot defend. It accrues silently with every undocumented decision and comes due in full the first time one is challenged.
The evidence layer is a pipeline, not a product. A question enters, and a defensible answer and a durable artifact leave. Each of the five components produces one output a challenger can inspect, and each closes one of the six failure modes from Section 02, all on an open-weight model the team runs itself.
| Failure mode | Closed by | Mechanism | Inspectable output |
|---|---|---|---|
| I. Unstated Assumptions | C1 Pre-Register | Claim and assumptions written and hashed before any evidence is gathered | Registered claim |
| II. Selective Evidence | C2 Investigate | Search runs to a stopping rule, every source recorded whether it agrees or not | Source ledger |
| III. No Falsification | C1 Pre-Register | A Kill Threshold set in advance states what would prove the answer wrong | Falsification criterion |
| IV. Buried Conflict | C3 Surface Conflict | Disconfirming evidence is reported in the answer, not discarded on the way | Conflict log |
| V. Unreproducible Path | C5 Provenance | The full reasoning trace is captured so a second run reaches the same place | Reproducible trace |
| VI. Absent Provenance | C5 Provenance | Every conclusion is linked to the evidence it depends on, retraction-aware | Evidence artifact |
The economics of defensible discovery invert the ones most organizations priced for. A hosted frontier model charges for every question, forever. A local open-weight model is paid for once, in hardware and setup, and answers the next question at a marginal cost that approaches zero. Published work shows a frontier model's quality can be matched at up to 98 percent lower cost in the best case [3]. Past a modest volume, local is simply cheaper.
But the inference bill is not the cost that matters. The cost that matters is being confidently wrong. The Defensibility Gap is paid not in cents per query but in the single decision a fluent, unexamined answer sends in the wrong direction, and that bill arrives whether the answer was hosted or local. Evidence is the only thing that lowers it.
Cost of inaction. A wrong decision made on an unexamined answer carries rework, delay, and the loss of credibility when it surfaces. None of it lands as an AI line item. All of it is Evidence Debt coming due.
Cost of control. The evidence layer is a six-to-twelve-week build for four to five people, on hardware already owned. It runs on the same local model the answers already use, so it adds rigor, not infrastructure.
Payback logic. One avoided wrong decision of consequence exceeds the build cost. The layer does not need to catch many. It needs to catch the one that would otherwise have shipped.
Each gate is a measurable exit condition, not a status meeting. A program that cannot make one workflow defensible in six weeks has no business extending the method across an estate. The pilot is deliberately small, because the discipline, not the scale, is what has to be proven first.
A reversible draft tolerates a fluent answer. An irreversible or externally visible decision demands the full layer, regardless of how good the model is.
Who will challenge this, an auditor, a regulator, a board, a peer? The more adversarial the audience, the more of the evidence layer the answer has to carry.
Can the question and its sources leave your environment? When they cannot, the model must be local, which the closing capability gap now makes practical.
Acted on once, or relied on for months? A durable conclusion needs provenance so that a later retraction can be traced to the decisions it touched.
The closing argument. A capable model is now a commodity, and a commodity confers no advantage. The organizations that pull ahead will be the ones whose answers can be trusted, challenged, and reproduced, on models they run themselves. That trust is what the evidence layer builds, and the fifteen weeks start whenever you do.
Answer honestly for your most consequential AI-assisted workflow. Count the YES answers, then read your level on the band below and the component gaps it points to. Designed to be completed in pen.
| No. | Tests | Question | Yes | No |
|---|---|---|---|---|
| 1 | C1 | Before a consequential answer, is the claim and its assumptions written down in advance? | ||
| 2 | C1 | Is a falsification criterion, what would prove the answer wrong, set before the work starts? | ||
| 3 | C2 | Does the investigation record every source it consulted, including the ones that disagreed? | ||
| 4 | C2 | Does the search run to a stopping rule rather than stopping the moment an answer appears? | ||
| 5 | C3 | Is disconfirming evidence reported in the final answer, not discarded along the way? | ||
| 6 | C4 | Does an adversarial reviewer challenge the answer before it is released? | ||
| 7 | C4 | Is every answer checked against its own falsification criterion before release? | ||
| 8 | C5 | Can a second person reproduce the answer from the recorded reasoning alone? | ||
| 9 | C5 | Is every conclusion linked to the evidence it depends on, so a retraction can be traced? | ||
| 10 | LOCAL | Does consequential work run on a model you control, so sensitive questions never leave? |
This paper is a conceptual contribution. The DDF taxonomy, the five-component evidence layer, the four-posture model, and the terms Defensibility Gap and Evidence Debt are original to this work. The pre-registration and adversarial-review primitives it builds on are instantiated in the research agent Zorp and cited where used. Holistic evaluation of language models established that quality is multidimensional and not captured by any single score [4]; the evidence layer extends that principle from measuring a model to defending an answer.
Quantitative claims follow one discipline throughout, made visible in every exhibit. Solid fills carry cited data. Hatched fills and dashed curves labeled directional exist to make the shape of an argument legible, not to report measurements. No statistic in this paper is attributed to a source that cannot be independently verified.
Arjun Jaggi and Aditya Karnam Gururaj Rao write and advise at the intersection of enterprise AI strategy, governance, and defensible discovery. Their published body of work, spanning original frameworks, concept papers, and books, is read by executive teams navigating AI adoption and is available in full, without a paywall, at arjunjaggi.com.