CMU MLNYUMay 21, 2026arXiv:2605.22612

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio, Fei Fang, Bryan Wilder

AI Summary

This paper argues that the gap between healthcare LLM benchmark performance and real-world deployment stems from unacknowledged assumptions about user interaction, specifically related to task completion and real-world outcomes. They categorize these assumptions and demonstrate this gap's existence through a retrospective analysis of a healthcare RCT. To bridge this gap, the authors propose BenchmarkCards for documenting assumptions and staged evaluation for systematic testing.

Key Contribution

Healthcare LLM benchmarks can be misleading because they fail to capture critical assumptions about how users interact with models and how those interactions translate to real-world outcomes.

Abstract

Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit assumptions about how users interact with models that cannot be surfaced from benchmarks alone. To make this precise, we propose a classification of assumptions into two categories: task, which can be tested from conversation data alone, and outcome, which requires outcome data and behavioral studies for testing. Critically, outcome assumptions depend on human behavior, something that even well-designed benchmarks cannot directly observe. To demonstrate the operationality of this framework, we retrospectively analyze a healthcare RCT as a case study and find that the gap naturally separates into task and outcome gaps of roughly equal size. To address this, we make two contributions: first, we propose BenchmarkCards, an artifact that documents assumptions, and second, we propose staged evaluation, a procedure that systematically tests assumptions and evaluates performance.

Constitutional AI & AI Ethics Eval Frameworks & Benchmarks Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

Related Papers