Override rates are outcome evidence. Dashboards are not.
A clinical AI tool can pass every technical review and still leave the only question a medical executive actually has unanswered: did patients do better? What Links 6 and 7 of the evidence chain mean inside a hospital.
Most clinical AI assessments stop at Link 3 of the evidence chain: validation. Sensitivity, specificity, AUROC on a retrospective cohort — necessary numbers, and not one of them tells a chief medical officer whether the tool improved a single decision on a single ward. Links 6 and 7 — Outcome and Decision Utility — exist because technical performance and clinical benefit are different claims, established by different evidence.
The first outcome signal is behavioral: what did clinicians actually do with the output? Acceptance and override rates, stratified by unit and by clinician seniority, are the cheapest honest evidence available. A model no clinician overrides is either excellent or ignored — and the difference is the whole question. A model overridden 60% of the time on one ward and 5% on another is telling you its Context of Use was drawn wrong, whatever the AUROC said.
Sensitivity and specificity are not properties of a model; they are choices about whose harm you tolerate. A sepsis alert tuned for sensitivity floods nurses with false alarms until the alerts get ignored — alarm fatigue is a patient-safety failure, not an inconvenience. The trade-off belongs in the Context of Use as an explicit, signed decision: this threshold, for this population, accepting this false-positive burden, reviewed on this schedule.
A model no clinician overrides is either excellent or ignored — and the difference is the whole question.
Prospective design is the discipline that makes outcome claims defensible. A silent-run period — the model scoring live patients while clinicians work unaided — establishes the baseline that retrospective validation cannot. Pre-registering the outcome question before go-live ('time-to-antibiotics for patients flagged at threshold X') prevents the quiet substitution of whatever metric happened to improve.
Harm review needs a cadence, not a trigger. Waiting for an incident report means the review happens after the harm. The alternative is scheduled: a standing review of overridden alerts that turned out correct, accepted alerts that turned out wrong, and every case where the model was silent and should not have been.
None of this requires machine-learning expertise from the clinical side. It requires the same instincts medicine already applies to a new drug or protocol: pre-specified endpoints, a control condition, monitored rollout, scheduled review. The evidence chain's last two links are clinical-trials thinking applied to deployed software — and a governance platform's job is to make sure those links get built, recorded, and revisited, not asserted once and forgotten.