Search papers, labs, and topics across Lattice.
This study introduces Auditable CT phenotyping (ACT), a method that leverages report-derived radiological observations to assess the validity of clinical phenotypes predicted from CT scans. By training on a large dataset of over 38,000 patients and evaluating against five vision-language baselines, ACT demonstrated superior performance in zero-shot annotation and linear probing across 221 phenotypes. Notably, the findings reveal that a limited number of observations dominate the model's predictions, suggesting that ACT can identify misleading correlations in clinical data while maintaining accuracy.
ACT uncovers that only 97 observations drive the predictions for 221 clinical phenotypes, revealing potential shortcuts in medical imaging interpretations.
Medical image foundation models can predict clinical phenotypes from computed tomography (CT), but strong performance leaves open whether they read disease-specific findings or shortcuts that correlate with the diagnosis. We tested this in 221 electronic-health-record (EHR) phenotypes using Auditable CT phenotyping (ACT), built on report-derived radiological observations. We trained ACT on 38,317 patients, mined 376,194 observations and evaluated it in 25,183 held-out patients. ACT exceeded five vision-language baselines on zero-shot annotation, and CT-CLIP across 221 phenotypes from unseen CT pulmonary angiography, both under zero-shot scoring (0.651 versus 0.572) and under linear probing (0.709 versus 0.662). Reading each probe exposes what accuracy conceals: only 97 observations occupy the 221 rank-1 positions, and one phrase describing aortic and coronary calcification ranks first for 20 phenotypes, including osteoporosis, urinary tract infection and major depressive disorder. Restricting the bank to clinician-specified evidence redirects those probes onto phenotype-related observations in 86 phenotypes at no accuracy cost (0.751 versus 0.741). Accurate CT-based EHR phenotyping can therefore rest on observations that are not valid evidence for the coded phenotype and that ACT can identify and intervene on.