Search papers, labs, and topics across Lattice.
This paper introduces G-CARL, a novel grounded, checklist-aligned reinforcement learning framework designed for Patient-oriented Medical Report Interpretation (PMRI), which aims to enhance the accessibility and accuracy of medical report explanations for patients. By integrating multi-source retrieval for claim verification with context-aware checklists, G-CARL effectively balances the dual objectives of factual accuracy and user-centered communication in medical contexts. Experimental results show that G-CARL significantly outperforms existing methods in quality, precision, and alignment with patient needs, as validated by clinician evaluations.
G-CARL not only improves the accuracy of medical report interpretations but also ensures they are tailored to patient queries, outperforming traditional methods in both factuality and user satisfaction.
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements. To bridge this gap, we introduce Patient-oriented Medical Report Interpretation (PMRI), a novel open-ended multimodal generation task that requires models to explain medical reports in accurate and accessible language based on a user's query and dialogue history. These two objectives differ fundamentally in their verifiability, yet remain tightly coupled, making them difficult to optimize jointly under conventional supervised fine-tuning and holistic reinforcement learning paradigms. To address this challenge, we propose G-CARL, a grounded, checklist-aligned reinforcement learning framework that combines multi-source retrieval for atomic claim verification with context-aware, instance-specific weighted checklists for response coverage, providing structured supervision for factuality, user-demand satisfaction, and expression quality without constraining response diversity. We further construct MMedReport, a real-world PMRI benchmark, along with a clinician-designed three-dimensional evaluation protocol. Extensive experiments demonstrate that G-CARL consistently outperforms existing post-training baselines in overall quality, claim-level precision, and checklist recall. Pairwise preference evaluation by clinicians further confirms that G-CARL produces interpretations that are more accurate and better aligned with patient needs.