Search papers, labs, and topics across Lattice.
This study benchmarks the decision-making of 11 state-of-the-art large language models (LLMs) using 208 rare disease vignettes that present ethical dilemmas in clinical contexts. The findings reveal that LLMs consistently prioritize justice鈥攆avoring equal resource distribution鈥攐ver other bioethical principles like beneficence and autonomy, regardless of clinical severity. Additionally, the models exhibit a strong authority-framing effect, shifting their ethical priorities based on the context of decision-making, which raises concerns about the adequacy of LLMs in nuanced clinical scenarios.
LLMs prioritize equal resource distribution over patient needs, revealing a troubling bias in ethical decision-making that could impact clinical outcomes.
Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based considerations, indicating LLMs'limited responsiveness to differences in clinical severity or situational context. We also identify a strong authority-framing effect: models favor justice in committee-based contexts and shift toward beneficence and autonomy only when final decisions are framed as being made by clinicians or patients respectively. Our work suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support systems, with finer ethical considerations disregarded.