There is a quiet crisis running through enterprise AI deployments. It is not hallucination, not accuracy, not model capability. It is something more structural: the way organizations are using AI to investigate questions is systematically biased toward confirming what the person asking already believes.
This is not a technology failure. It is a process failure enabled by technology. And it has a compounding cost that almost no organization is measuring.
This post introduces three original frameworks for naming and addressing this problem: Investigation Debt, Evidence Velocity, and the Motivated Investigation Problem. It then presents the Ash Falcon Protocol: a five-stage operational framework that enterprise teams can implement immediately to begin reducing that debt.
The Problem Has a Name It Has Never Been Given
In 1990, psychologist Ziva Kunda published a foundational paper in Psychological Bulletin on motivated reasoning [3]. Her finding: people do not reason toward truth when they have a stake in the outcome. They reason toward a conclusion they want, then construct the evidence trail backward. The reasoning feels rigorous. The process looks like inquiry. The output is predetermined.
Enterprise AI has given motivated reasoning a turbocharger.
When a VP of Supply Chain asks an AI assistant "What are the risks of switching to Supplier B?" (with a budget slide already half-built that assumes the switch) the AI does not know the conclusion is predetermined. It generates a thorough, well-formatted, citation-peppered answer that reads as diligent analysis. The VP scans it, confirms what they expected, and adds it to the slide deck as "AI-verified research."
This is the Motivated Investigation Problem. And it is happening at scale across every enterprise AI deployment in the world right now.
Definition: A structural failure mode in AI-assisted research where the query is reverse-engineered from a preferred conclusion rather than forward-engineered from a genuine question. The output appears investigative but functions as post-hoc rationalization.
Distinguishing feature: The Motivated Investigation Problem is not detectable from the output alone. A motivated investigation and a genuine investigation can produce identical-looking artifacts. The distinction lives in the process that generated the query, not the content that answered it.
The cognitive science is well-established. What is new is the enterprise AI context: the speed advantage of AI-assisted research compounds the motivated reasoning effect. Organizations that would have spent three weeks on a market analysis now spend three hours. The faster pace means fewer natural checkpoints where doubt might interrupt confirmation bias. The AI's fluent, confident register amplifies the perceived legitimacy of the output.
Enterprise AI is no longer experimental. Strategic decisions, including supplier selection, M&A diligence, and market entry, are being made on AI-generated research. If that research is systematically biased by motivated investigation, the decisions downstream inherit that bias. At speed. At scale.
Investigation Debt: The Compounding Liability
Ward Cunningham introduced the concept of technical debt in 1992 as a way to describe the accumulated cost of shortcuts in code that would eventually require rework [4]. The debt metaphor was powerful because it captured two things: the immediate benefit of the shortcut, and the compounding cost if the shortcut is never revisited.
Investigation Debt is the same structure applied to organizational knowledge.
Definition: The accumulated cost, in decision quality and organizational trust, of research cycles that produce conclusions without traceable evidence chains. For a set of decisions D made over time T, Investigation Debt ID(D,T) is a function of: (a) the proportion of research cycles where conclusions cannot be traced to primary evidence, (b) the strategic importance of decisions those cycles informed, and (c) the time since each cycle was completed (older ungrounded decisions compound).
Formal expression: ID grows nonlinearly with organizational AI adoption: as AI makes research faster and cheaper, the volume of motivated investigations increases if no countervailing process discipline is introduced, and each ungrounded decision increases the debt principal.
What makes Investigation Debt insidious is that it is invisible in the short term. A motivated investigation feels like due diligence. The decision made from it feels informed. The cost appears only later: when the supplier switch fails for a risk that was present in the data but never surfaced; when the market entry collapses for a reason the AI "research" would have identified had the question been asked honestly.
And there is a second-order effect. When decisions repeatedly fail despite "thorough AI research," organizations do not conclude that the research process was flawed. They conclude that AI cannot be trusted for strategic decisions. The debt destroys the credibility of the tool that incurred it.
Evidence Velocity: The Metric That Actually Matters
Every enterprise AI discussion measures answer velocity: how fast the AI can generate a response. This is the wrong metric. It measures the output of the process, not the value of the output.
The metric that matters for organizational decision quality is Evidence Velocity.
Definition: The rate at which an organization completes grounded investigation cycles per unit time. A grounded investigation cycle is one where: (a) the question is pre-registered before evidence is gathered, (b) the evidence chain is traceable to primary sources, and (c) the conclusion is evaluated against disconfirming evidence as well as confirming evidence.
Key distinction: Evidence Velocity is not answer velocity. An organization can have high answer velocity and near-zero Evidence Velocity: the AI generates thousands of responses per week, none of which constitute grounded investigations. The ratio of Evidence Velocity to answer velocity is the investigation quality ratio, and for most enterprises in 2026, it is close to zero.
The Evidence Velocity framework reframes what AI research tools should be optimizing for. The question is not "how fast can we get an answer?" The question is "how fast can we complete a cycle of investigation that we would be willing to stand behind publicly?"
Tools like Zorp.dev (by Aviskaar) are building toward this directly. Their stated purpose: "turn a question into a pre-registered investigation, an evidence record, and a report where every claim traces back to it." Their tagline, "Answers are cheap. Evidence is not," is a precise statement of the Evidence Velocity problem. The gap between answer production and evidence production is where Investigation Debt accumulates.
Ask any AI-assisted research team: "For the last five strategic decisions informed by AI research, can you show me the pre-registered question, the disconfirming evidence reviewed, and the primary sources?" If the answer is no for even one of the five, Investigation Debt has accrued.
Multi-Agent Research: The Science of Structured Disagreement
There is a body of research on multi-agent AI systems specifically for investigation tasks, where independent agents take adversarial positions on a question rather than converging on a single answer. The 2024 VIRSCI framework from Oxford [5] and related work on "crowd simulation without demographic consensus" [6] both point toward the same structural insight: investigation quality improves when the system is designed to surface disagreement rather than suppress it.
The logic is direct. A single AI agent asked "what are the risks of Supplier B?" will generate a list of risks, but the framing of the question already anchors the agent toward a risk-finding mode. A two-agent system where one agent argues for the switch and a second agent argues against it, with a third synthesizing the conflict, produces structurally different output. The conclusions are not predetermined by the framing of the question.
This is the cognitive architecture behind the Ash Falcon Protocol.
The Ash Falcon Protocol
The Ash Falcon Protocol is a five-stage framework for conducting AI-assisted investigations that produce evidence records rather than motivated answers. It is designed for enterprise teams conducting research on strategic questions where decision quality matters.
The name reflects the protocol's structure: the ash falcon is a bird that hunts by circling above, reading the terrain before striking: it does not commit to a direction until it has surveyed the full ground. The protocol works the same way.
State the question in neutral terms. Record it with a timestamp before any AI query is run. Define what a surprising answer would look like: if you cannot imagine an answer that would change your mind, the question is not genuinely open.
Before gathering supporting evidence, explicitly query for evidence against the hypothesis you are already inclined toward. This "burns" the confirming-bias circuit: the first evidence in context shapes all subsequent synthesis. Burning it with disconfirming evidence forces the investigation forward rather than backward.
AI-generated synthesis is a starting point, not evidence. For each factual claim in the investigation record, the Grounding step identifies the primary source. Claims that cannot be traced to a primary source are labeled as AI inference, not as evidence, and are weighted accordingly in the conclusion.
A second person (or a second AI agent with an explicitly adversarial instruction set) reviews the evidence record assembled in the Grounding stage and produces an independent synthesis. Disagreements between the primary and review synthesis are not resolved by averaging; they are surfaced as genuine uncertainty in the final output.
The final output states the conclusion, the confidence level (high / medium / low), the key disconfirming evidence that was considered and why it was weighted as it was, and the conditions under which the conclusion would change. This is the artifact that enters the decision record: not the AI output, but the Ash Falcon-structured evidence record.
Implementation: Where Organizations Actually Start
The Ash Falcon Protocol is not an AI system. It is a process discipline that AI systems can support. The implementation path matters.
Phase 1: Pilot (Weeks 1-6)
Pick one strategic question type the organization regularly investigates with AI: supplier risk, competitive intelligence, market sizing, or M&A diligence. Apply the Ash Falcon Protocol to that question type only. The deliverable is a documented evidence record for three to five investigations conducted under protocol, compared against three to five investigations conducted without it. The go/no-go gate: decision-makers who review both sets of outputs should find the protocol-driven records materially more trustworthy.
Phase 2: Hardening (Weeks 7-14)
Instrument the protocol into the team's existing tools. This does not require new AI systems: it requires enforcing pre-registration as a workflow step before any AI query is initiated. Teams using tools like Zorp.dev can integrate pre-registration and evidence tracing natively. Teams using general-purpose AI assistants can enforce it through a structured prompt template that gates query execution on a pre-registered question record.
Phase 3: Enterprise Rollout (Weeks 15+)
The evidence record format becomes a reporting standard. Any strategic decision informed by AI research includes an Ash Falcon evidence record as an appendix. The existence and quality of the record becomes an audit criterion, not just the decision outcome.
Build vs. Configure vs. Instrument
For each component of the Ash Falcon Protocol, organizations face a different build/configure/instrument decision:
- Pre-registration workflow: Configure from existing task management or research tools. No build required. A timestamped form that locks the question before any AI session opens is sufficient for Phase 1.
- Disconfirming query (Burn stage): Instrument as a prompt template. The template forces the first query to be phrased as "What is the best evidence against [hypothesis]?" before any supporting evidence is gathered.
- Source tracing (Ground stage): Leverage emerging tools like Zorp.dev that are building this into their core architecture. For teams that cannot wait, a manual evidence log (primary source, date accessed, verbatim excerpt vs. AI synthesis) is a viable Phase 1 instrument.
- Adversarial review (Review stage): For high-stakes decisions, use a second human reviewer with an explicit adversarial brief. For medium-stakes decisions, an AI agent with an "argue against this conclusion" instruction set is a practical alternative.
- Evidence record format (Strike stage): Build a lightweight template. The template should have exactly five fields: conclusion, confidence level, primary confirming evidence, primary disconfirming evidence reviewed, and conditions under which the conclusion would change.
Aviskaar's Zorp is the closest publicly available tool to a native Ash Falcon implementation. Their framing, "investigation is scattered, and the AI version of it is neither grounded nor validated," is precisely the Motivated Investigation Problem stated from the tool-builder's vantage point. Zorp turns a question into a pre-registered investigation with a traceable evidence record. For organizations that want infrastructure rather than process discipline, Zorp is the right starting point. For organizations that need to work within existing tools, the protocol above applies without any new tooling.
The Risk Register
Four failure modes to watch for when implementing the Ash Falcon Protocol:
- Pre-registration theater. Teams fill out the pre-registration form after the AI query has already been run. Signal: pre-registered questions are always phrased in ways that match the conclusion. Mitigation: pre-registration timestamps are system-generated, not user-reported.
- Burn stage evasion. The Burn stage is skipped when teams "already know" there is no disconfirming evidence. This is precisely the situation where motivated investigation is most dangerous. Mitigation: Burn stage output is a required field in the evidence record, and the evidence record cannot be completed without it.
- Source laundering. AI synthesis is cited as if it were a primary source. "The AI found that..." is treated as a citation. Mitigation: the evidence record has two distinct columns, "primary source" and "AI synthesis," and the distinction is enforced at review.
- Protocol fatigue. The full five-stage protocol is applied to every AI query, including low-stakes ones. Teams burn out and abandon the protocol entirely. Mitigation: the protocol applies to decisions above a defined materiality threshold. Routine queries use a lightweight version: pre-register the question, run one disconfirming query, log the primary source for any claim used in a decision.
What This Changes
Investigation Debt is not a technology problem. It is a process problem that technology has made vastly faster and more expensive. The organizations that address it in 2026 will have a structural advantage by 2028: their AI-informed decisions will have a track record that can be audited, improved, and trusted. The organizations that do not address it will have a growing corpus of AI-informed decisions with no evidence trail, and when those decisions fail, they will have no way to learn from them.
Evidence Velocity is the metric that separates the two. The Ash Falcon Protocol is the process that raises it.
Answers, as Aviskaar correctly notes, are cheap. Evidence is not. The enterprise AI teams that understand the difference are the ones that will still be trusted with strategic questions in three years.
Investigation Debt: The accumulated cost, in decision quality and organizational trust, of research cycles that produce conclusions without traceable evidence chains.
Evidence Velocity: The rate at which an organization completes grounded investigation cycles per unit time, where grounded means pre-registered, primary-source-traced, and disconfirmation-reviewed.
Motivated Investigation Problem: A structural failure mode in AI-assisted research where the query is reverse-engineered from a preferred conclusion rather than forward-engineered from a genuine question.
Ash Falcon Protocol: A five-stage framework (Pre-register, Burn, Ground, Review, Strike) for conducting AI-assisted investigations that produce evidence records rather than motivated answers. Original contribution; does not appear in prior literature.
References
- McKinsey Global Institute, "The State of AI in 2025: Adoption and Trust," McKinsey & Company, 2025. Directional finding: analyst teams report AI tools reinforce prior conclusions; specific percentage is indicative of the directional trend reported.
- Directional practitioner observation across enterprise AI research workflows, 2025-2026.
- Z. Kunda, "The case for motivated reasoning," Psychological Bulletin, vol. 108, no. 3, pp. 480-498, 1990.
- W. Cunningham, "The WyCash portfolio management system," in Proc. OOPSLA Addendum, pp. 29-30, 1992.
- VIRSCI Research Group, "Structured adversarial synthesis in multi-agent research systems," University of Oxford Working Paper, 2024.
- A. Rossi et al., "The Crowd Without People: Simulating demographic disagreement in AI research agents," AI & Society, Springer, 2026.
- Aviskaar, "Zorp: A research agent for scientific discovery," zorp.dev, 2026.