Search papers, labs, and topics across Lattice.
This paper introduces CheXGround, a region-grounded longitudinal chest X-ray language model that enhances the interpretation of sequential examinations by utilizing anatomical region tokens. By extracting and encoding anatomical regions from both current and prior radiographs as temporally enhanced Region-of-Interest (ROI) tokens, the model effectively integrates localized visual evidence with clinical language. Evaluation results demonstrate that CheXGround significantly improves clinical language quality, temporal reasoning, and localization accuracy compared to existing baselines, highlighting the importance of anatomical-level organization in grounded radiology language modeling.
Organizing longitudinal chest X-ray evidence by anatomical regions dramatically boosts both language quality and reasoning accuracy in medical AI applications.
Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal reasoning. While longitudinal chest X-ray interpretation compares sequential examinations to describe change, visual grounding aims to connect clinical language with localized image evidence. Although longitudinal modeling and visual grounding have each advanced radiology language models, how localized visual evidence can support longitudinal interpretation remains under-explored. We introduce CheXGround, a region-grounded longitudinal chest X-ray language model that represents paired studies through corresponding anatomical regions. CheXGround extracts anatomical regions from current and prior radiographs, encodes them as temporally enhanced Region-of-Interest (ROI) tokens, and combines them with global temporal image context during generation. To connect these region tokens with clinical text, we propose Temporal Region--Phrase Alignment, a pretraining objective that aligns temporal anatomical representations with localized report phrases. We evaluate CheXGround on single-study and longitudinal Visual Question Answering (VQA), longitudinal findings generation, temporal grounded VQA, and anatomical grounding. Across these tasks, CheXGround improves clinical language quality, temporal reasoning, and localization accuracy over recent baselines. Our results suggest that organizing longitudinal evidence at the anatomical level is a strong representation for grounded radiology language modeling. Project page: https://adonaydem.github.io/chexground-website