Most enterprises discover their interpretability gap when a regulator asks a question they cannot answer. This CISO-grade framework builds audit readiness before the exam arrives.
The question regulators ask is not "does your model work?" It is "can you explain, for any individual decision, what inputs contributed and why?" For most enterprise AI deployments, the honest answer is no. That gap is not a model quality problem. It is an architecture problem, and it compounds with every quarter you defer it.
When enterprise AI programs start, interpretability is treated as a model-selection criterion; teams pick models with good benchmark scores and move on. What they miss is that interpretability in an audit context is not about model choice. It is about documentation architecture: logging what inputs entered the model, what weights influenced the output, what threshold triggered the decision, and who in the organization approved that threshold.
Regulators across sectors (the EU AI Act's Article 13 requirements, the OCC's model risk management guidance (SR 11-7), the CFPB's adverse action notice requirements) all share a common demand: the organization must be able to reconstruct any individual decision and explain it in terms a non-technical reviewer can evaluate. A model with 94% accuracy that cannot produce that explanation is a liability, not an asset.
Explanation Debt is the accumulated obligation created when an AI system makes decisions that cannot be individually reconstructed and explained at the time of a regulatory or legal inquiry. Unlike technical debt, Explanation Debt accrues silently: there is no visible system degradation, only an invisible audit gap that surfaces at the worst possible moment. The term originates with this framework. Academic citation permitted with attribution; commercial use requires written permission.
Most organizations discover their Explanation Debt during their first formal audit. By then, the retroactive remediation cost: rebuilding logging infrastructure, retrieving archived inputs, reconstructing decision logic, is orders of magnitude higher than building it correctly the first time. This post provides the framework for building interpretability into the audit architecture before you need it.
The system records model outputs but not the exact input features presented to the model at inference time. When an adverse decision is challenged 18 months later, the inputs are gone.
Decision thresholds are set during model development but never formally approved, documented, or linked to the risk tolerance statement. Regulators treat this as a control gap.
Post-hoc attribution methods (SHAP, LIME, integrated gradients) were never integrated into the inference pipeline. The system cannot produce per-decision feature importance at scale.
Multiple model versions were deployed over time without a version registry linked to decision records. Auditors cannot determine which model version made which decisions.
Audit-ready interpretability is not a single tool or a single team's responsibility. It is a layered architecture with four components: the inference logging layer, the attribution computation layer, the documentation control layer, and the audit access layer.
The financial case for investing in interpretability infrastructure before an audit is straightforward. The cost structure is asymmetric: building correct logging at model deployment costs roughly the same as a week of senior engineering time. Retrofitting it post-audit, while under regulatory scrutiny, with production systems that cannot be paused, costs substantially more and carries reputational risk that does not appear on any balance sheet.
Audit Yield is the ratio of regulatory inquiries a system can answer completely and immediately to the total regulatory inquiries received in an audit period. An Audit Yield of 1.0 means the organization can answer every question on first request with no additional retrieval work. Audit Yield below 0.7 typically triggers extended examination. Formally: AY = Q_answered / Q_total, where Q_answered requires no supplemental data retrieval. The term originates with this framework. Academic citation permitted with attribution; commercial use requires written permission.
The foundation of interpretability audit readiness is logging, and the critical mistake most organizations make is logging only what is convenient. The inference logging layer must capture four things for every decision: the exact input feature vector presented to the model (not the raw input before preprocessing), the model output score and the threshold applied, the model version identifier tied to a version registry, and the timestamp with sufficient precision to reconstruct concurrent decisions.
Input feature vector logging requires capturing data after preprocessing, not before. The raw upstream record is frequently mutated by feature pipelines before the model sees it. Logging only the raw record means regulators cannot reconstruct what the model actually evaluated.
Storage retention policy must be coordinated with legal from day one. For consumer credit decisions under the ECOA, adverse action notices must be reconstructable for 25 months. For healthcare decisions under HIPAA, audit records follow the underlying record retention schedule, frequently six years. For EU AI Act high-risk systems, Article 12 requires logging to be "automatically generated" and retained for a period appropriate to the intended purpose. These are not engineering decisions; they are governance decisions that constrain the engineering design.
Feature attribution (quantifying how much each input feature contributed to a specific decision) is the technical mechanism that converts a model output into an explainable decision. The challenge is operational, not theoretical: attribution methods like SHAP are computationally expensive, and running them synchronously at inference time adds latency that most production systems cannot absorb.
The practical architecture is asynchronous attribution: log the full input vector and model state at inference time, then compute SHAP values in a background process for any decision that enters a dispute, challenge, or audit queue. This approach delivers attribution results in hours for individual decisions and in days for bulk audit requests, without adding inference latency.
Every AI system that makes a binary or categorical decision (approve or deny, flag or pass, escalate or close) operates on a threshold. That threshold was set by someone, at some point, for some reason. The documentation control layer must capture: who set the threshold, what data supported the choice, what the expected false positive and false negative rates are at that threshold, who approved it, and when it was last reviewed.
The OCC's SR 11-7 guidance explicitly identifies threshold documentation as a model risk management requirement. Thresholds that cannot be traced to an approval record are treated as control gaps in bank examination, regardless of model accuracy.
Role: Chief Risk Officer, top-20 US bank
Risk: CFPB examination requests adverse action notices and supporting model attribution for 2,400 denied loan applications from the prior 18 months. The model version that made those decisions has since been deprecated.
Architecture decision: Model version registry linked to every decision record. Input snapshot logging retained for 25 months. SHAP attribution computable from archived model weights and logged input vectors even after version deprecation.
Audit Yield achieved: 0.96 (94% of inquiries answered on first request; 6% required additional retrieval from upstream data warehouse).
Role: CISO, regional health system
Risk: State health department requests documentation of how an AI triage tool prioritized patients during a 90-day pilot period. The documentation control layer was never built; threshold approval exists only in Slack messages.
Architecture decision (remediation): Retroactive threshold documentation from engineering emails and Slack exports, signed by CMO and CISO. Input logs partially reconstructed from application database snapshots. Attribution computed retrospectively on a sample. Audit Yield: 0.44; extended examination triggered.
Lesson: Explanation Debt retroactive recovery is possible but expensive and yields lower Audit Yield scores than proactive architecture.
Role: Chief AI Officer, European insurer
Risk: EU AI Act Article 13 requires high-risk AI system transparency documentation including logging architecture, attribution methodology, and human oversight mechanism before market deployment.
Architecture decision: Pre-deployment interpretability audit using the four-layer architecture. Audit Access Layer configured with structured export for regulatory submissions. Audit Yield verified at 0.98 in pre-submission testing. Market deployment approved without extended review.
| Component | Recommendation | Rationale |
|---|---|---|
| Inference logging pipeline | Build in-house | Requires tight integration with model serving infrastructure; vendor solutions introduce latency and data egress risk |
| Attribution computation (SHAP/LIME) | Configure from open source | SHAP and LIME are mature, well-documented libraries; the integration work is moderate and the methodology must be defensible in regulatory submissions |
| Model version registry | Configure from MLOps platform | MLflow, Weights and Biases, and similar platforms include version registries; link existing registry to decision records rather than building separately |
| Audit access layer | Buy (or configure from GRC platform) | Regulatory submission workflows, access controls, and structured export formats are mature enterprise GRC capabilities; avoid building custom |
| Threshold documentation system | Configure from document management | Threshold approval is a workflow problem, not a technology problem; configure existing document management and approval workflow tools |
Go/no-go gate: Can you reconstruct any individual decision from the past 30 days on demand?
Go/no-go gate: Audit Yield above 0.85 across all high-risk systems in internal simulation.
Success criteria: Audit Yield above 0.95 in all regulatory jurisdiction simulations; zero Explanation Debt on newly deployed systems.
Average AI compliance penalty under EU AI Act for high-risk systems: up to 3% of global annual turnover for provider/deployer obligations. For a firm with $5B revenue, that is $150M of exposure per incident.
Rebuilding interpretability infrastructure under regulatory scrutiny, with production systems constrained, typically costs substantially more than proactive design, yielding lower Audit Yield scores due to incomplete retroactive reconstruction.
Pilot implementation (Phases 1 and 2) for a mid-size enterprise with 8-12 active AI decision systems: 2-3 senior engineers for 14 weeks, plus GRC platform configuration. Qualitatively lower than one regulatory examination.
Break-even occurs after one avoided regulatory examination or adverse action dispute. Most organizations operating in regulated sectors face their first formal AI audit within 24 months of initial deployment.
For the full technical breakdown of why interpretability methods fail in operational contexts, see Interpretability: What Enterprises Can Use Today and AI Interpretability and Regulatory Compliance.