Search papers, labs, and topics across Lattice.
This paper introduces MMed-Bench-IR, a comprehensive benchmark for evaluating multilingual medical information retrieval across six languages and three distinct tasks: cross-lingual medical QA retrieval, concept discrimination, and multilingual evidence retrieval. By ensuring no overlap in concepts or queries among tasks, the benchmark provides a clearer assessment of system capabilities in real-world multilingual clinical contexts. The evaluation of ten systems reveals significant performance disparities, highlighting that existing English-only benchmarks fail to capture critical cross-lingual deficiencies, with nDCG@10 scores plummeting from 0.818 in English to 0.056 in Japanese.
Multilingual medical retrieval systems face a staggering performance drop, with scores plummeting from 0.818 in English to just 0.056 in Japanese, exposing a critical gap in current benchmarks.
Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora. Multilingual medical retrieval demands three capabilities: cross-lingual alignment, concept discrimination, and evidence retrieval. However, existing benchmarks evaluate these only in isolation, leaving the interaction between biomedical expertise and multilingual coverage unmeasured. We introduce MMed-Bench-IR, a benchmark designed to disentangle these axes across 6 languages and three structurally heterogeneous tasks: (1) cross-lingual medical QA retrieval with 6,127 queries grounded in the Unified Medical Language System (UMLS), (2) concept discrimination over 4,975 confusion sets at three difficulty tiers, and (3) multilingual evidence retrieval for RAG with 2,040 quality-assured queries. The three tasks share zero concept and query overlap by design, ensuring that aggregate scores reflect genuine capability breadth. Evaluation of ten systems across six paradigm families reveals severe cross-lingual failure: biomedical encoders that score 0.818 nDCG@10 in English drop to 0.056 in Japanese, a gap that English-only benchmarks cannot detect.