Search papers, labs, and topics across Lattice.
This study benchmarks ten frozen 3D CT encoders on their ability to detect a variety of findings in thoracic CT scans, utilizing methods such as $k$-nearest neighbors and linear probing. The results reveal that while models leveraging fine-grained image tokenization and vision-language alignment tend to excel, a lightweight supervised encoder can rival larger models, indicating that explicit labels can effectively compensate for model scale. Notably, the research highlights that the detectability of findings is primarily influenced by their contrast against surrounding tissue and spatial extent, with small, low-contrast lesions remaining a significant challenge across all models tested.
Detectability of findings in 3D CT scans hinges more on physical characteristics than model architecture, revealing a critical bottleneck in diagnostic performance.
Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing generalizable representations of anatomy and pathology. To evaluate their diagnostic breadth, we benchmark ten frozen CT encoders across three cohorts of thoracic CT scans, including an unseen internal clinical dataset, using $k$-nearest neighbors, zero-shot prompting, and linear probing. We find no universal state-of-the-art, with rankings fluctuating significantly depending on the evaluation context. While models combining fine-grained image tokenization with vision-language alignment generally perform best, a lightweight supervised encoder remains highly competitive, demonstrating that explicit labels can effectively substitute for scale. Crucially, rather than model architecture, we observe that the primary determinant of performance is a physical bottleneck: a finding's detectability scales with its contrast against surrounding tissue and its spatial extent. Through controlled within-organ comparisons, we empirically demonstrate that widespread or high-contrast abnormalities, such as devices and effusions, are reliably recovered. Conversely, small, low-contrast focal lesions remain a persistent challenge across all evaluated encoders. We attribute this to the inherent limitations of globally pooled embeddings, suggesting that accurately representing small, low-contrast structures will require region- or lesion-level pretraining.