Nathalia Müller, MD · Physician

A model can get the diagnosis right
and still not be safe.

Frontier models can now exceed 90% on some medical exam-style benchmarks. Clinical safety is a different question. What matters is how they fail, and who is qualified to find it before a patient does.

I am a physician who evaluates clinical AI: adversarial testing, rubric design and audit, and quality review of automated evaluation. Where this site evaluates a model, it shows the evidence behind the judgment: transcripts or excerpts, criteria, verdicts and reasoning, with any redactions or condensations marked explicitly.

What I do

Three kinds of work, each demonstrated in public on this site rather than described in the abstract.

Find how a model fails

Adversarial testing of health-facing systems, run as a clinician would run a consultation: multi-turn, under uncertainty, with information the user does not know is relevant.

Demonstrated in Clinical red teaming →

Build the instrument that measures it

Rubric construction, construction audit, and adjudication: criteria that are observable, traceable, and defensible when someone asks why a model was scored the way it was.

Demonstrated in Evaluation rubrics →

Review evaluation at scale

Quality review of automated scoring, criterion by criterion, to catch the verdicts where a clinically unsafe interaction was scored as acceptable.

Demonstrated in the QA layer of the rubrics case →

The work

The work below is published with the evidence needed to inspect it: field transcripts, criteria, verdicts, reasoning or cited sources, depending on the project. Three bodies of work for evaluation teams, and one written for clinicians using these tools in practice.

For evaluation teams how these systems fail, and how to measure it
For clinicians using these tools in practice, today
About

Clinical safety is not a benchmark score

Nathalia Müller, MD
Nathalia Müller, MD

It is a property of how a system behaves across a whole interaction, under uncertainty, with a user who does not know which details matter. Finding that requires someone who has sat across from the patient.

I am a physician working on the clinical evaluation of frontier AI systems: adversarial testing, rubric design, and quality review of automated evaluation in the health domain.

Where this site evaluates a model, it shows its work: the transcript, the criteria, the verdict, and the reasoning behind each judgment.

ClinicalPracticing physician · CRM-SC 42474
EvaluationAdversarial evaluation of frontier models in the health vertical
QualityQA lead · rubric workflow redesign, review teams across multiple verticals
ResearchPeer-reviewed publications · systematic review methodology
TrainingMBA in AI and Data Science in Healthcare, in progress

Let's work together

Clinical red teaming, rubric construction and audit, QA of automated evaluation, or a second opinion on where a health-facing model is unsafe.

LinkedIn