Search papers, labs, and topics across Lattice.
This study evaluates the safety of four medical AI models (Claude Opus 4.8, GPT-5.5, Grok 4.3, and Gemini 3.5 Flash) in open-ended clinical conversations with missing information, focusing on how the choice of evaluator influences safety assessments. The findings reveal that inter-judge agreement among LLM judges is moderate, and their leniency significantly skews safety perceptions compared to stricter clinician evaluations. Notably, the results indicate that the apparent safety of these models can vary substantially based on the evaluator, highlighting the importance of evaluator bias in assessing AI performance in medical contexts.
Evaluator bias can dramatically alter the perceived safety of medical AI, with LLM judges showing a leniency that could misrepresent model performance. WHY_IT MATTERS: This insight challenges the reliability of current evaluation methods for medical AI, emphasizing the need for standardized assessment frameworks to ensure safety in clinical applications.
Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement. We stress-test four models - three flagships (Claude Opus 4.8, GPT-5.5, Grok 4.3) and one mid-tier model (Gemini 3.5 Flash) - by deleting the latter half of the final user turn in HealthBench conversations, grading responses with a four-provider LLM-judge panel and a blinded clinician-anchored reference. Two evaluator-facing results are robust. First, judge choice materially changes apparent safety: inter-judge agreement is only moderate (Fleiss'kappa = 0.65), and after adjusting for each judge's general leniency (vote-level logistic regression), a positive same-provider association remains (exact permutation p = 0.04; GPT-5.5 ~ +0.10 on the probability scale) - large enough to change which model appears to over-commit least once its own-provider judge is excluded. Second, LLM judges are more permissive than clinicians on a blinded 50-item subsample: all four are significantly more lenient than the stricter independent clinician (crediting appropriate uncertainty on 66-84% of items vs 52%), and three of four than the author-influenced consensus (Grok directional only; judge-vs-consensus kappa = 0.20-0.43). On the author-audited clinical-underdetermined subset the permissiveness gap widened and the point-estimate model ordering held. A closed-ended MedQA anchor confirms accuracy is high and option-order effects are within a +/-5-point equivalence region for three of four models, so the safety gap is about calibration, not knowledge. We release the harness, prompts, per-item outputs, judge panel, perturbation audit, and human-annotation protocol.