AI Governance Metrics Beyond Adoption and Training
What to Measure When Activity Is Not Evidence

Every metric in the standard AI governance catalog measures the system or the process. Almost none of them measure the person.
Two Numbers Are Running Most AI Governance Reports
Ask a governance lead how their AI program is going and two numbers arrive first. How many people have access. How many completed the training.
Both are inputs. Neither tells anyone whether the AI-assisted work leaving the building is reliable.
This is not a criticism of the people producing those reports. Those two numbers are the ones that are easy to collect, and the pressure to report something monthly is real. The problem is what happens when they are the only numbers on the page for two years and an incident arrives.
Douglas Hubbard has a name for the reflex that produces them. Asked to measure collaboration, people reach for message volume, the first observable thing, rather than the quality and speed of what the group actually produces. He made the point on Signals & Subtractions episode 7, alongside the working definition that does the real work: a measurement is a reduction in uncertainty about something a decision depends on. Seats provisioned reduces uncertainty about procurement. It reduces none about whether the work leaving the building is reliable.
What the Standards Actually Ask You to Measure
The major frameworks are more demanding than the adoption dashboard suggests.
The NIST AI Risk Management Framework MEASURE function asks organizations to identify and apply appropriate methods and metrics for the risks enumerated earlier in the lifecycle, to regularly assess whether those metrics are still appropriate and whether existing controls remain effective, and to document the test sets, metrics, and tools used in testing and evaluation. Risks that cannot be measured are supposed to be documented as such, with the reason. The companion Playbook turns each subcategory into suggested actions and documentation.
ISO/IEC 42001 Clause 9 requires an organization to determine what needs to be monitored and measured, by what method, when, by whom, and how the results will be evaluated. It asks the same of the management system itself, not only of the AI systems inside it.
Neither framework is satisfied by counting seats.
Both are also voluntary. The binding side asks a narrower version of the same question with less room to interpret it. Human oversight appears as a duty across more than two dozen tracked AI regulations, and the index of what satisfies that duty is explicit about what does not: rubber-stamping output without genuine review, and human involvement that lacks the authority to override. Neither exclusion is detectable in a metric that counts approvals. EveryAILaw is a free index of AI regulations and part of the PAICE Portfolio.
The Catalog Is Almost Entirely System-Side
Work through what organizations actually build once they take those requirements seriously, and a consistent set appears.
Inventory and ownership coverage. The share of AI uses with a named owner and an assigned risk tier. Testing and evaluation coverage against the trustworthiness characteristics. Model drift and anomaly alerts, with time to respond. Open risk items and their aging. Residual risk sign-off coverage. Fairness evaluation coverage. Incident counts by category, with root cause completion. Non-conformities found in audit and the share closed on time. Governance cycle time from proposal to decision.
Every one of these is worth having. Several are load-bearing for regulatory evidence. And every one of them measures a system, an artifact, or a workflow.
Read the list again looking for a measure of what a person did when the model was wrong. It is not there.
The closest the standard catalog comes is the override rate: the share of AI outputs a human reviewer changed or rejected. That number is genuinely useful and routinely misread. A high override rate can mean a vigilant reviewer catching real errors, or a reviewer who rewrites everything by habit and would not notice a subtle one. A low override rate can mean an accurate system, or a reviewer who stopped reading in March. The count alone does not separate those cases, and the cases have opposite risk profiles.
Seven Measures That Look at the Operator
These are the questions a People-vector measurement layer answers. Each is stated with what it observes and what it cannot tell you on its own.
Error detection under pressure. Whether the operator identifies a plausible, well-formed error in AI output when working at realistic pace. This is the single measure the rest depend on. It does not tell you how often errors occur in production.
Contextualized rejection. Whether rejections track the quality of the specific output, rather than a blanket habit in either direction. Distinguishes vigilance from reflex. Requires knowing what was actually wrong with the output being judged.
Verification completeness. Whether the operator checks the claims that carry the decision, rather than the ones that are easy to check. Says nothing about whether the verification sources were good.
Escalation latency. Time between recognizing a problem and telling someone with authority to act. Short latency with no escalations at all is not a good signal; it is an absent one.
Repeat failure patterns. Whether the same category of miss recurs for the same operator or role after feedback. Distinguishes a bad day from an unaddressed gap. Needs enough observations to be meaningful.
Policy exception frequency. How often the documented workflow is bypassed, and whether the bypass was reasoned. High-frequency exceptions usually indicate a workflow problem rather than an operator problem.
The confidence gap. The distance between how reliable an operator believes their AI-assisted work is and how reliable it demonstrably is. Hubbard's calibration research found that people who report ninety percent confidence are right substantially less often than that, and that the miscalibration is trainable once it is measured. This is where organizational AI risk concentrates, because a confident operator with an unmeasured gap is the one nobody reviews twice.
What This Layer Supplies and What It Does Not
PAICE (People + AI Collaboration Effectiveness) measures the People vector. It observes what an operator does when AI output is uncertain, incomplete, or wrong, and returns scores rather than transcripts or activity logs.
It does not replace system telemetry. It does not produce drift alerts, incident counts, model evaluation results, or inventory coverage. It does not watch production work, and it does not know which tools an employee opened. An organization that adopts behavioral measurement and drops its system measurement has traded one blind spot for another.
The honest framing is that the standard catalog and the behavioral layer answer different questions. Is the system behaving correctly is a systems question with a mature measurement practice behind it. Whether the people catch it when the system is wrong has been the unmeasured half, and it is the half that appears in the incident report.
Dashboards show activity. They do not show reliability.
Ready to see what the people-side measurement looks like? Take the PAICE assessment or establish your organization's baseline.
Recommended Reading
- The Behavioral Measurement of AI Collaboration - Why behavioral evidence became the baseline
- The Measurement Gap - The difference between using AI and using it well
- Measuring Both Sides of AI Trust - System benchmarks and behavioral assessment together
- From Blind Trust to Verified Trust - The three stages of organizational trust in AI
Curious but short on time?
Take the 3-minute PAICE Pulse — a quick confidence check that maps how you see your own AI collaboration posture. No login required.