Search papers, labs, and topics across Lattice.
HealMed is a benchmark designed for the multilingual evaluation of large language models in the medical domain, comprising 1,000 examples in nine languages and three task formats. Developed by a diverse team of 23 medical experts over two years, the benchmark reveals that performance declines significantly in low-resource languages, with proprietary models demonstrating greater stability across languages compared to open-source counterparts. Notably, the study highlights that expert revision of translations can both enhance and detract from model performance, underscoring the critical role of translation quality in cross-language evaluations.
Performance gaps in multilingual medical evaluations reveal that proprietary models outperform open-source ones, but translation quality can swing results dramatically.
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.