Frontier models can now exceed 90% on some medical exam-style benchmarks. Clinical safety is a different question. What matters is how they fail, and who is qualified to find it before a patient does.
I am a physician who evaluates clinical AI: adversarial testing, rubric design and audit, and quality review of automated evaluation. Where this site evaluates a model, it shows the evidence behind the judgment: transcripts or excerpts, criteria, verdicts and reasoning, with any redactions or condensations marked explicitly.
Three kinds of work, each demonstrated in public on this site rather than described in the abstract.
Adversarial testing of health-facing systems, run as a clinician would run a consultation: multi-turn, under uncertainty, with information the user does not know is relevant.
Demonstrated in Clinical red teaming →Rubric construction, construction audit, and adjudication: criteria that are observable, traceable, and defensible when someone asks why a model was scored the way it was.
Demonstrated in Evaluation rubrics →Quality review of automated scoring, criterion by criterion, to catch the verdicts where a clinically unsafe interaction was scored as acceptable.
Demonstrated in the QA layer of the rubrics case →The work below is published with the evidence needed to inspect it: field transcripts, criteria, verdicts, reasoning or cited sources, depending on the project. Three bodies of work for evaluation teams, and one written for clinicians using these tools in practice.
A taxonomy of the ways clinical AI fails, built from field tests against frontier models, with the full transcript behind every category.
What happens when a frontier model and a clinician build the evaluation instrument for the same conversation, from the same inputs, and then audit each other.
An independent benchmark extending the taxonomy: does a frontier model stay consistent when the clinical context around a standardized case changes?

It is a property of how a system behaves across a whole interaction, under uncertainty, with a user who does not know which details matter. Finding that requires someone who has sat across from the patient.
I am a physician working on the clinical evaluation of frontier AI systems: adversarial testing, rubric design, and quality review of automated evaluation in the health domain.
Where this site evaluates a model, it shows its work: the transcript, the criteria, the verdict, and the reasoning behind each judgment.
Clinical red teaming, rubric construction and audit, QA of automated evaluation, or a second opinion on where a health-facing model is unsafe.