Search papers, labs, and topics across Lattice.
This study systematically evaluates seven intensity normalization methods for a 3D U-Net model tasked with meniscus segmentation from knee MRI, addressing the critical issue of model generalizability across varying imaging protocols. The methods, including standard scaling, histogram-based techniques, and a Gaussian Mixture Model (GMM), were assessed on both internal and external datasets, revealing that while Z-score, Ny'ul histogram matching, and CLAHE performed better, the overall impact of normalization was overshadowed by the significant performance drop due to domain shifts. These findings underscore the necessity for additional strategies beyond normalization to ensure robust clinical deployment of deep learning models in medical imaging.
Intensity normalization can enhance MRI segmentation performance, but its benefits are minimal compared to the challenges posed by domain shifts.
Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but underexplored factor affecting model generalisability is intensity normalisation, particularly for magnetic resonance imaging (MRI), where image intensities vary across scanners and protocols. In this study, we systematically compared seven normalisation methods and their impact on the performance of a 3D U-Net model for meniscus segmentation from knee MRI. The methods included standard scaling approaches, histogram-based techniques, and a Gaussian Mixture Model (GMM)-based method. Models were trained on the IWOAI 2019 dataset and evaluated on both internal and external test sets (SKM-TEA) to assess generalisability. Performance was similar internally but differences were significant on external data, with Z-score, Ny\'ul histogram matching, and CLAHE showing greater robustness than other methods. However, these differences were small compared to the significant performance drop observed between datasets. Overall, while intensity normalisation had a measurable effect on model generalisability, its impact was limited relative to the effects of domain shift, highlighting the need for complementary strategies for robust deployment.