Enterprise AI  ·  Evaluation & Governance

You Can't Govern What You Can't Test.
Enterprise AI Has No Tests.

Your organization has deployed AI systems that make behavioral claims. None of them have been tested. No eval was written before the first prompt shipped. This is the governance gap nobody is naming: Eval-Driven Development, and why every enterprise AI program needs it before the next model update.

Arjun Jaggi  ·  August 24, 2026  ·  14 min read
0
AI governance controls requiring a runnable eval suite before go-live [1]
#1
behavioral testing identified as the primary gap in enterprise LLM governance [3]
Art. 9
Regulatory article requiring documented testing measures for high-risk AI systems [2]

The Problem Belongs to the CTO

Every enterprise that would never ship a line of software without a test suite is shipping AI systems with zero behavioral specifications. The same organization that requires 80% code coverage in CI/CD has deployed an LLM-powered customer escalation system with no documented behavioral claims and no eval to verify them. When a model update arrives, nobody knows what changed because nobody ever specified what it was supposed to do.

This is not a quality assurance problem. It is a governance problem. You cannot audit a system whose behavior was never specified. You cannot demonstrate regulatory compliance through a behavior that was never tested. You cannot confidently update a model when you have no mechanism to detect whether the update changed anything. Eval-Driven Development is the practice that closes this gap: write the evaluations before you ship the system, not after something breaks.

The CTO and Chief AI Officer own this problem. The cost of not addressing it is not visible until a model update silently changes a behavior that was carrying a compliance control, a customer promise, or a risk decision, and nobody finds out until a regulator, a customer, or a competitor does.

Original Term: coined here

Dark Behavior is any output, decision, or response pattern of a deployed AI system that was never formally specified, never tested with a runnable evaluation, and cannot be audited against a documented behavioral claim. Dark Behavior is not a bug: the system may be performing exactly as it was informally intended to. The risk is that it is performing without proof, without version history, and without any mechanism to detect when that behavior changes. In a high-risk AI deployment under In any deployment, it is an ungoverned business decision running at scale.

Original Term: coined here

Eval Surface is the total set of behavioral claims about a deployed AI system that are backed by runnable, versioned evaluations. An organization's Eval Surface defines how much of its AI system's behavior it can actually prove, audit, and monitor. The complement of the Eval Surface is Dark Behavior. Most enterprise AI systems today have an Eval Surface that covers fewer than a third of the behavioral claims their stakeholders believe the system makes, based on structural inference from how AI systems are built and deployed without formal specification disciplines. Organizations that have achieved Eval-Driven Development operate with an Eval Surface that expands with every new system capability, ensuring behavioral claims and the tests that verify them grow together.

Why Evals Are Not a Quality Step: They Are the Specification

The practitioner community has been circling this insight for several years. The clearest version of it: if you cannot write an eval for a behavior, you do not actually know what that behavior is. A prompt that instructs a model to "respond professionally and helpfully" is not a behavioral specification. It is an aspiration. A behavioral specification says: for this input class, the system must always return a response that satisfies these measurable criteria, and here is a runnable test that verifies it.

This reframes the question from "how do we measure quality after deployment" to "how do we define what we're building before we build it." The CheckList framework from Ribeiro et al. [3] demonstrated in 2020 that behavioral testing of NLP models: testing capability types, not just aggregate accuracy, surfaces failure modes that benchmark metrics consistently miss. That insight applies with even greater force to enterprise AI systems, where the behavioral claims are not academic but operational: the system must not recommend competitor products, must escalate certain keywords to compliance queues, must refuse to answer out-of-scope questions. These are testable claims. They are almost never tested.

Key Insight

The behavioral contract between an AI system and its enterprise context (what the system will always do, what it will never do, and what it must do consistently under edge conditions) lives in the prompt. The same Prompt Debt dynamic documented in The Prompt Is Business Logic means that when the prompt is ungoverned, the behavioral contract is ungoverned. An eval suite is the executable form of that contract.

The Eval-Driven Development Architecture

Fig. 01: Eval-Driven Development Pipeline: From Behavioral Specification to Governed Deployment
Prompt Library Compliance Rules Edge Case Library Model Updates
  • [3] M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh, "Beyond Accuracy: Behavioral Testing of NLP Models with CheckList," Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), arXiv:2005.04118.
  • [4] P. Liang et al., "Holistic Evaluation of Language Models," arXiv:2211.09110, 2022.
  • [5] E. Perez et al., "Red Teaming Language Models with Language Models," arXiv:2202.03286, 2022.
  • Excited about AI, innovation, and growth?

    Start a conversation