Search papers, labs, and topics across Lattice.
The authors scale 3D radiological foundation models by training complementary convolutional and transformer architectures on 2.1 million multi-modal (CT, MRI, PET) volumes across 125 datasets. Evaluated across 108 diverse downstream tasks, the models substantially outperform existing 3D baselines while revealing that generalist performance in 3D vision is split: ConvNets consistently dominate spatially localized tasks like segmentation and detection, whereas transformers excel at global semantic reasoning and frozen-feature transfer. By integrating post-hoc dataset-aware topological alignment into standardized pipelines like nnU-Net, the framework offers an immediately deployable standard for volumetric medical representation learning.
No single foundation model rules 3D radiology: across 2.1 million volumetric scans and 108 benchmarks, ConvNets consistently outperform transformers at dense spatial localization, while transformers decisively win on global semantic reasoning.
Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a single pretrained model can support diverse downstream tasks. Here we present nnFoundation, complementary convolutional and transformer-based 3D radiological foundation models. Developed within the Human Radiome Project (THRP), nnFoundation is trained on 2.1 million CT, MRI, and PET image volumes from 125 institutional and public datasets. We evaluate them across 108 tasks spanning segmentation, detection, classification, report generation, and image retrieval, including evaluations under domain shift, by external partners and in low-data and low-compute regimes. Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging. However, performance follows a consistent task-dependent structure: the convolutional nnFoundation model dominates spatially localized tasks, whereas the transformer-based nnFoundation model excels in tasks requiring global semantic reasoning and in frozen-feature settings. Dynamically aligning the foundation model topology with the dataset characteristics post-hoc further improves transfer across heterogeneous 3D settings. These results show that transferable 3D radiological performance is governed not by a single universal model, but by the interplay of scalable pretraining, complementary architectures, and dataset-aware adaptation. We release nnFoundation models integrated into nnU-Net and nnDetection, enabling immediate application across established radiology workflows.