Search papers, labs, and topics across Lattice.
The authors investigated the tension between cross-lingual consistency and cultural adaptation in multilingual medical LLMs by surveying 356 experts across NLP, medicine, and anthropology in the US, Germany, and Spain. They found that while standard benchmarks penalize cross-lingual variance as error, human experts are fundamentally divided鈥攁nthropologists favor adaptation, while clinicians split along geographic lines鈥攁nd persona-prompted LLMs fail to replicate this nuanced divergence by heavily overestimating consistency preferences. These findings challenge the core assumption of multilingual medical NLP benchmarks that medically correct outputs should remain invariant across languages.
Multilingual medical benchmarks penalize cross-lingual answer variation as model error, but real-world clinicians are sharply split on whether AI should enforce universal consistency or adapt to local cultural contexts.
Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.