Validation is not a certificate. It's a boundary.
A validation report answers the question it was designed around. The harder question is what happens at the edge of it.
Scientific AI can perform well on every benchmark and still lack the evidence that it is trustworthy for the specific decision being made. The model is measured; the decision is not. That gap is where expensive failures live.
Consider the standard sequence. A method is described as 'human-relevant' or 'predictive.' A validation study is commissioned. The study compares the method against the legacy approach it was built to replace — because that is where the historical data sits. The results are strong. A certificate-shaped artifact is produced, circulated, and filed.
Then the method meets a decision: a candidate progresses or it does not; a patient is eligible or is not; a submission is filed or delayed. And the certificate goes quiet, because it was never about this decision. It named no population, no operating conditions, no excluded indications. It cannot say where confidence stops.
The certificate goes quiet, because it was never about this decision.
Validation does not certify a technology. It defines the boundary inside which a decision-maker can act. A boundary has coordinates: who decides what, for which population or test articles, under which conditions, against which comparator, and — critically — where the method must not be used. Written that way, a claim becomes something a study can confirm or refute, and something a reviewer can evaluate.
The practical test is simple to state: could an independent reviewer infer the decision, the consequence of error, and the required evidence from the current Context of Use? If not, a strong validation package assembled before the decision was defined may still be inadequate — and it is better to learn that before the filing, not after.