Search papers, labs, and topics across Lattice.
This paper introduces MEGA-CDP, a novel benchmark designed to evaluate the adherence of large language models (LLMs) to clinical decision pathways (CDPs) defined by clinical practice guidelines. By constructing a dataset from 2,274 guidelines and generating 42,353 clinical cases, the authors assess LLM performance in both single-turn and multi-turn settings, revealing significant challenges in reliable clinical decision support. The findings underscore the necessity for CDP-oriented evaluation frameworks, highlighting the limitations of current models in adhering to established medical guidelines.
Current medical LLMs struggle to consistently follow clinical guidelines, revealing a critical gap in their decision-making capabilities.
Following clinical decision pathways (CDPs) defined by clinical practice guidelines is essential for safe and reliable medical decision-making. However, existing medical large language model (LLM) benchmarks mainly evaluate final-answer accuracy, providing limited evaluation of models'ability to adhere to guidelines. To address this gap, we introduce MEGA-CDP, a benchmark for evaluating whether medical LLMs can generate guideline-adherent CDPs using provided guidelines as references. MEGA-CDP is constructed from 2,274 English and Chinese clinical practice guidelines through a guideline-to-case pipeline, yielding 42,353 clinical cases with explicit reference CDPs. It supports both single-turn vignette and multi-turn interactive settings, and introduces a CDP-oriented evaluation framework for measuring pathway consistency. Experiments on 16 representative LLMs show that reliable clinical decision support remains challenging for current models, demonstrating the need for CDP-oriented evaluation and the value of MEGA-CDP for advancing guideline adherence in medical LLMs.