Search papers, labs, and topics across Lattice.
This study critically evaluates the performance of large language models (LLMs) on Polish medical exams, revealing that traditional multiple-choice question answering (MCQA) methods may inflate perceived clinical competence due to inherent biases and guessing strategies. By introducing a more rigorous benchmark with over 15,000 questions and structural modifications, the authors demonstrate a significant drop in performance for the top model, Qwen3.5-122B, by 28.4 and 31 percentage points on English and Polish exams, respectively. The findings underscore the importance of evaluation design in accurately assessing LLM capabilities in medical contexts, highlighting that standard MCQA scores may not reliably indicate true medical competence.
LLMs can appear competent in medical contexts, but a rigorous evaluation reveals a staggering drop in performance that questions their true clinical abilities.
Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical ability due to guessing strategies and answer biases. To address these limitations, we introduce an expanded and more challenging benchmark based on Polish medical exams, adding over 15,000 questions, two new domains, and four structural modifications that reduce MCQA-specific artifacts and better test reasoning. We evaluate 21 LLMs and show that evaluation design strongly affects results. Under our harder setup, the best model (Qwen3.5-122B) drops by 28.4 and 31 pp on English and Polish exams, respectively. Despite low evidence of data contamination, standard MCQA scores do not reliably reflect true medical competence. To facilitate further research, we make our benchmark publicly available.