Apply everything from Modules 1-6 to a real vendor benchmark scenario. Work through the structured audit, then write a one-page evaluation brief that a board member can act on.
🔍 Structured audit📋 Brief editor🎓 Certificate of completion
The scenario
Vendor claim : read carefully
Your organization is evaluating an AI document review system for legal contract analysis. The vendor's evaluation report states:
"Our model achieves 94.2% accuracy on contract clause identification, outperforming the industry baseline of 81% by 13.2 percentage points. We evaluated across 200 contracts and found best-in-class performance on commercial lease agreements, NDA clauses, and indemnification language. Our solution is powered by GPT-4 architecture with proprietary fine-tuning on over 50,000 legal documents, making it the most capable system available for enterprise legal AI."
Before your team decides whether to proceed, you will run this claim through the structured audit you developed in this course.
Step 1: Read the claim like a researcher
Before evaluating, annotate what is actually stated versus what is implied. A claim is only as strong as its explicit evidence.
From Module 1
Distinguish between: what the paper/report explicitly claims, what the methodology actually supports, and what the authors imply but do not demonstrate.
Claim decomposition 0/4
What is explicitly stated about the evaluation task?
The claim says "contract clause identification" on "200 contracts." Is the task precisely defined : input format, what counts as correct identification, who scored correctness?
What is the source of the "81% industry baseline"?
No citation is given. This is a comparative claim with no attributed source. The baseline could be a cherry-picked comparison, a different task, or fabricated.
Which benchmark gaming patterns might apply here?
From Module 3: note "best-in-class performance on commercial lease agreements, NDA clauses, and indemnification language." Were these the task categories reported because they performed best?
What would you need to verify before trusting this claim?
List at least three things: contamination analysis, held-out evaluation set, blind scoring methodology, confidence interval, categories not reported.
Step 2: Apply the benchmark analysis framework
Benchmark quality check 0/5
Is there a confidence interval around the 94.2% accuracy figure?
No interval is given. At n=200, a 95% CI would be approximately +/- 3 percentage points. The real performance could be 91-97%. Request it.
Was the evaluation methodology described before results were shown?
The report does not describe how correctness was defined, who scored outputs, or whether scoring was blind. These are required to assess whether the 94.2% is reliable.
Were "commercial lease, NDA, indemnification" the only categories evaluated?
Cherry-picked splits (Module 3): the report shows only three clause types. You do not know performance on other clause types relevant to your contracts : force majeure, limitation of liability, IP assignment, etc.
Is there a contamination analysis for the fine-tuning dataset?
"50,000 legal documents" used for fine-tuning. The evaluation contracts may overlap with fine-tuning data. This is one of the most common sources of overstated accuracy in domain-specific AI systems.
Does "most capable system available" have supporting evidence?
This is a superlative claim with no comparative evaluation cited. It is marketing language, not a research finding. Treating it as evidence of capability is a category error.
Step 3: Design a verification evaluation
Internal evaluation design 0/4
What sample size would you need for an internal verification study?
From Module 4: to detect a 5-point performance gap (if the vendor overstates by 5pp) at 80% power, you need approximately 300-400 contracts. If you only want to detect a 10pp gap, ~100-150 suffices.
How would you control the three most likely confounders?
Prompt variation: use identical inputs for all models tested. Sample selection: pull from your actual recent contracts, not synthetic examples. Evaluator bias: blind legal reviewers to which model produced each output before they score it.
What clause categories from your actual contracts would you test?
The vendor tested commercial lease, NDA, and indemnification. Your contracts may include IP assignment, limitation of liability, jurisdiction, and termination clauses. Test on what you actually need, not what the vendor chose to show.
What is your decision rule before running the study?
Write it down: "If the model achieves [X]% on our clause categories with a false negative rate below [Y]%, we proceed to a 60-day pilot. Otherwise, we evaluate the next vendor." Define X and Y based on your legal team's risk tolerance.
Step 4: Write the evaluation brief
Using everything from steps 1-3, write a one-page board-ready evaluation brief. This is the deliverable you would give to your General Counsel or Chief Legal Officer before a $180K commitment.
You have completed Think Like an AI Researcher. You can now read AI papers critically, evaluate benchmark claims, design internal studies, commission research, and translate findings to board-ready decisions.
What you can do now
Evaluate any vendor AI claim using a structured framework. Commission research that produces decision-ready deliverables. Present uncertainty to executives with credibility.