Search papers, labs, and topics across Lattice.
This paper introduces MIRA-Ev, a novel benchmark for evaluating clinical NLP systems through granular evidence detection and relational reasoning, addressing the limitations of traditional multiple-choice question answering (MCQA) that overlooks the nuances of clinical reasoning. By re-annotating Spanish MIR licensing-exam cases with detailed span-level premises and claims, the authors create a comprehensive resource that includes directed support and attack relations, enhancing the evaluation of model performance. The benchmark is significant as it is released in Spanish, English, and Basque, marking the first clinical argumentation resource in Basque and providing a structured framework for assessing argument quality in clinical contexts.
MIRA-Ev reveals that traditional MCQA methods fail to capture the complexities of clinical reasoning, highlighting the need for a more nuanced evaluation framework.
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument mining benchmark built on Spanish M\'edico Interno Residente (MIR) licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and released in parallel Spanish (native), English, and Basque versions, the first clinical argumentation resource in Basque. MIRA-Ev organizes evaluation into a three-tier task hierarchy: evidence sentence retrieval, argumentative component extraction, and relation classification.