Search papers, labs, and topics across Lattice.
This paper introduces a Polish-language medical visual question answering (VQA) benchmark derived from Board Certification Examination questions, evaluating various vision-language models on their ability to utilize visual evidence. Despite the challenging nature of the task, the best-performing model achieves only 79.0% accuracy, with most models underperforming compared to human benchmarks, particularly on image-dominant questions. The analysis reveals that models rely more on textual information than visual cues, indicating a significant gap in their ability to effectively integrate visual evidence in medical contexts.
Vision-language models struggle to leverage visual evidence in medical VQA, with only one model surpassing human performance on a subset of questions.
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0\% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.