Search papers, labs, and topics across Lattice.
This study investigates the efficacy of Monte Carlo (MC) Dropout for voxel-level uncertainty estimation in glioma segmentation from multiparametric MRI, highlighting its potential to identify segmentation errors in clinically critical sub-regions. Through an empirical analysis of two models on 126 BraTS21 patients, the authors found that while MC Dropout maintained high segmentation accuracy, it also revealed significant discrepancies in uncertainty calibration, particularly in the UNet-Res model. The results emphasize that standard metrics like Dice scores can obscure critical failures in model reliability, necessitating region-specific calibration assessments for clinical applications.
Despite high uncertainty-error alignment, a model can still miscalibrate confidence in critical regions, posing serious patient safety risks.
Glioma segmentation in multiparametric MRI is a critical component of treatment planning. A segmentation model that fails silently on treatment-critical sub-regions represents a patient safety risk that overlap-based metrics such as Dice scores cannot expose. We ask whether voxel-level uncertainty estimation via Monte Carlo (MC) Dropout can reliably identify segmentation errors in clinically critical sub-regions, and whether calibration failure modes are detectable from standard reporting metrics alone. In an empirical two-model case study on 126 BraTS21 patients, we evaluate a high-performance pretrained SegResNet and a locally trained UNet with residual units (UNet-Res). MC dropout preserved segmentation accuracy ($|螖\text{Dice}|$ $<0.01$) while achieving strong uncertainty-error alignment (AUROC for entropy (H) $\approx$0.97), indicating uncertainty correctly ranks erroneous voxels above correct ones. Entropy-based patient stratification identified a high-uncertainty subgroup with substantially lower segmentation performance (median whole-tumour Dice $0.835$ vs. $0.925$), supporting uncertainty as a practical triage signal. However, global alignment can mask important region-specific differences. Despite similar AUROC, UNet-Res exhibited near-zero enhancing tumour entropy ($0.054$) and Expected Calibration Error (ECE) of $0.915$, with a Dice of only $0.714$, indicating severely miscalibrated confidence on the most clinically critical sub-region, a failure mode invisible to standard Dice and AUROC reporting. These findings demonstrate that strong uncertainty-error alignment is necessary but insufficient for clinical safety: sub-region-specific calibration assessment must accompany AUROC evaluation when selecting models for clinical deployment.