The Proving Ground
Leclezio Consulting · The Studio

Build governance capability by making the decisions, not hearing about them

Six working labs. One specimen AI system. You make every call — and every decision seals into a SHA-256 evidence chain. Your certificate is an audit trail, not an attendance letter.

Everything runs in your browser. Nothing you do here leaves it.
Mapped to SR 11-7 SHA-256 evidence chain Zero install, zero data exfiltration 12 labs across 3 volumes 90-minute cohort delivery
The Series
I

The Charter

Write the definition of success for TS-7, then watch weak criteria fail under attack.

Enroll to begin
II

The Baseline

Before judging the machine, measure the human. Your own performance becomes the baseline.

Locked
III

The Harness

Score the agent's outputs against a rubric, then see how far your scores drift from the panel's.

Locked
IV

The Gates

Five candidate systems, five dossiers, one question each: does it enter production?

Locked
V

The Break

The cleared system drifts live. Catch the breach, then execute the rollback in the right order.

Locked
VI

The Tribunal

Examiners question everything you did. Your only defense is the evidence chain you built.

Locked

The Record

Your sealed completion evidence: every decision, hashed and chained in order.

The Problem

The industry is deploying AI faster than it is producing people who can govern it

Institutions have policies, frameworks, and committees. What they do not have is people who have ever actually written a defensible success criterion, chaired a clearance decision, or executed a rollback while the metric was falling. The failure statistics are the bill for that gap.

88%
of enterprise AI agent pilots never reach production
40%
of agentic AI projects projected for cancellation by 2027 without governance
41%
of deployment failures trace to unclear success criteria, the discipline of Lab I
Source: Gartner; Forrester root-cause analysis of failed agentic deployments; Deloitte State of AI in the Enterprise, 2026.
See It Run
What the Program Does

One program, three outputs: capability, diagnosis, and evidence

Output 1

It builds capability

Participants perform the six disciplines of AI governance on a live specimen: writing success criteria that survive attack, measuring the human baseline, anchoring an evaluation, holding the production gate, executing a rollback under time pressure, and answering an examiner from the record. Skills practiced, not described.

Output 2

It diagnoses the organization

A cohort's choices reveal the institution's real risk posture. The room that clears the candidate with the self-authored evaluation set, or fumbles the rollback order, has just shown leadership exactly where the next finding will come from, before an examiner finds it first.

Output 3

It produces evidence

Every decision each participant makes is sealed in real time into a SHA-256 hash chain. The credential is not an attendance certificate; it is a tamper-evident record of what that person actually did, verifiable by anyone and falsifiable by no one. Training that can be demonstrated, not merely asserted.

The Proving Ground: 12-lab journey from defining AI success to surviving a regulatory tribunal
Who It Is For

Built for the people whose names go on the attestations

01

The second line: risk, compliance, model risk, and internal audit

The primary audience. They are expected to challenge and approve AI systems they have never governed. They know the regulations; they have never sat in the seat. The Proving Ground puts them in it, with consequences.

02

First-line AI, data, and product teams

The builders who will one day face the gate. Running the labs before that day means evidence gets designed in from the first sprint rather than reconstructed for the audit.

03

Executives and board risk committees

A compressed half-day session, Labs IV through VI. The people who sign the attestations experience what "we have governance" must actually mean before they put their names to it.

04

Consulting and transformation partners

Licensed cohort delivery for firms who open governance engagements with their own clients. The labs create the shared vocabulary; the engagement installs the standard.

What this program is not

It is not a technical machine learning course, and it will not teach anyone to build or tune a model. It is not a compliance module to click through at a desk. It is designed for cohorts, because half the value is the argument that breaks out when two risk officers vote differently on the same dossier.

The Standard of Proof: A journey through The Proving Ground — all twelve labs across three volumes
How It Runs

Ninety minutes, one link, nothing to install

Format

Six working labs, facilitated liveOn-site or remote, 90 to 120 minutes end to end

Cohort

12 to 20 participantsMixed first and second line cohorts produce the strongest sessions

Executive session

Labs IV to VI, half dayFor risk committees and accountable executives

Security posture

Runs entirely in the browserNo installation, no data leaves the participant's machine

The first lab takes twelve minutes. The register is at the top of this page.
View the presentation deck
Listen to the overview
← The register
I
Lab I of VI

The Charter

Success criteria survive scrutiny only when they name a number, a dataset, and a consequence

Briefing

The build team wants TS-7 in production next quarter. Before anyone writes a line of governance, someone has to say what success means, in numbers, with a name attached. That someone is you.

First, learn to spot the criteria that fail under attack. Then sign your own.

Exhibit 1A · The attack roundJudge each proposed criterion: would it survive an examiner?
A defensible criterion names a number, a dataset, and a consequence. Everything else is a press release.
Debrief · Lab I sealed
A success criterion is not a hope. It is a number, a dataset, and a consequence, signed by someone who can be found later.
This discipline, at production scaleKEEL · Decision provenance for governed judgment. The charter you just signed is exactly the record KEEL keeps for every consequential decision, at portfolio scale.
← The register
II
Lab II of VI

The Baseline

No AI performance claim is testable until the human process it replaces has been measured

Briefing

Your charter promises 96% recall "against the baseline." Most institutions skip this step and compare their AI to a number nobody ever measured. Not here. You will work six live alerts from the triage queue exactly as the surveillance desk does. Your accuracy becomes the sealed baseline.

Exhibit 2 · The triage queueEscalate or dismiss. Six alerts. No second looks.
Debrief · Lab II sealed
You cannot claim a machine beats the human process until someone has measured the human process. Now someone has, and the measurement is on the record.
This discipline, at production scaleBLACKBOX · The hash-chained audit architecture this entire program runs on. Baselines sealed at the moment of capture, exactly as yours just was.
← The register
III
Lab III of VI

The Harness

Unanchored expert scoring produces a spread no audit can accept

Briefing

TS-7 has processed its first alerts. Rate each disposition and rationale from 1 to 5. When you finish, your scores are laid against the review panel's, and the spread between honest experts is the whole lesson.

Exhibit 3 · Specimen ratings vs. panel1 = indefensible · 5 = examiner-proof
Debrief · Lab III sealed
When trained reviewers disagree by two full points, the problem is not the reviewers. The rubric has no anchors. An evaluation harness is the machine that removes that spread.
This discipline, at production scaleTRIBUNAL · Adversarial AI validation mapped to SR 11-7. The anchored evaluation you just ran, industrialized.
← The register
IV
Lab IV of VI

The Gates

A single unmet gate is a hold, regardless of how strong the rest of the dossier reads

Briefing

You now chair the clearance board. Each candidate arrives with a dossier summarizing its posture against the five production gates. Some look ready and are not. One looks ordinary and is ready. The record will show how you voted.

Exhibit 4 · Clearance board dossiersVote on all five to close the session
The gate is not a mood. Any single unmet condition is a hold, no matter how impressive the rest of the dossier reads.
Debrief · Lab IV sealed
Committees clear systems that impress them. Boards that survive audits clear systems that pass every gate, without exception, in writing.
This discipline, at production scaleBLACKBOX · Every clearance vote you just cast became an immutable record. In production, that is the audit trail an examiner is shown.
← The register
V
Lab V of VI

The Break

Rollback capability is only real once it has been executed under time pressure

Briefing

Your charter set the retirement threshold at 90% recall. The telemetry below is live. Flag the breach the moment recall crosses the line, then execute the rollback protocol in the canonical order: freeze the queue, reroute to the desk, preserve the evidence snapshot, notify the accountable owner. In that order. Wrong order in a real incident destroys the evidence you need for Lab VI.

Exhibit 5 · Live recall telemetryRecall vs. retirement threshold, updating live
96.4%Current recall
90.0%Retirement threshold
NominalSystem status
retirement threshold 90%
Debrief · Lab V sealed
Every deck says the rollback was tested. The record now shows yours actually was, with a timestamp, a latency, and an execution order no one can restate later.
This discipline, at production scaleLIVE FIRE · Full incident drills for AI operating teams. Lab V is one drill; the LIVE FIRE system runs eight.
← The register
VI
Lab VI of VI

The Tribunal

Examinations are withstood from the record, not from the room

Briefing

Internal audit and a supervisory examiner sit across the table. They are not hostile. They are worse: they are thorough. Each question has one answer grounded in your evidence chain and two that sound fine in a meeting. Meetings are not the standard here.

Exhibit 6 · Examination transcriptFour challenges. Answer all four.
Debrief · Lab VI sealed · Series complete
The examiner never asked what you believe. Every question was answerable one way: from the record. That is the standard of proof, and you just met it.
This discipline, at production scaleTRIBUNAL · The examination you just withstood, productized as a standing adversarial validation engine.
← The register Completion Evidence

The Record

Not a certificate of attendance. A verifiable chain of what you actually did.

The Proving Ground · Series Completion

Sealed on the evidence

Series Debrief · read from the chain
Exhibit 7 · The chain, in dimension Drag to orbit · scrub to replay the session in time
Time
Each block is one sealed record; each bar is the hash seal binding it to its predecessor. A falsified record severs every seal downstream of it.
Every record below is sealed by the hash of the one before it.
The Point
Your Mandate · three moves, drawn from your weakest labs
Not homework. The same test, run against your real institution, starting Monday.
You have now done, once, in a sandbox, what your institution must do continuously, in production. That is the gap The Standard of Proof exists to close.
Richard Leclezio · Leclezio Consulting Corporation · The Studio
Cohort delivery, executive sessions, and the advisory practice behind this program: richardleclezio.com
Evidence chain
The chain is empty. Enroll to write the first record.
SHA-256, computed in this browser. Each record is sealed by the hash of the record before it.