NYUStony BrookUniversity of Hawai'i at M ānoaUniversity of Hawai'i Cancer CenterMay 6, 2026arXiv:2605.05082

External Validation of Deep Learning Models for BI-RADS Breast Density Prediction from Ultrasound Images

Yuxuan Chen, Arianna Bunnell, Yanqi Xu, Haoyan Yang, Thomas K. Wolfgruber, John A. Shepherd, Yiqiu Shen

AI Summary

This paper externally validates three deep learning models (DenseNet121, ViT-B/32, ResNet50) for predicting BI-RADS breast density from ultrasound images on a new cohort of 2000 exams, including cancer cases and matched controls. The models achieved high AUROC scores for extremely dense (0.868-0.899), fatty (0.814-0.838), and scattered density (0.764-0.799) breasts, but lower performance for heterogeneously dense breasts (0.699-0.729), with DenseNet121 showing the best overall performance (micro-averaged AUROC 0.885). Incorporating AI-derived density into a risk prediction model did not outperform using mammography-reported density, suggesting areas for further improvement.

Key Contribution

Turns out, deep learning models trained to predict breast density from ultrasound images generalize surprisingly well to external datasets, but still struggle with heterogeneously dense breasts.

Abstract

We externally validated three deep learning models (DenseNet121, ViT-B/32, and ResNet50) for predicting mammographic breast density from breast ultrasound exams on an independent cohort. The external validation set comprised 2,000 ultrasound exams, including 500 cancer cases defined by an initial negative exam (BI-RADS 1 or 2) followed by a cancer diagnosis within 6 months to 10 years, and 1,500 negative controls matched by manufacturer and study year. Performance was measured using patient-level AUROC across four density categories: A (fatty), B (scattered), C (heterogeneous), and D (extremely dense). As a downstream assessment, we also evaluated 10-year risk prediction by incorporating age and AI-derived density into the Tyrer-Cuzick model and comparing performance against a reference model using age and mammography-reported density. All three models performed best in extremely dense breasts (AUROC 0.868-0.899), with strong performance in fatty (0.814-0.838) and scattered density (0.764-0.799), and lower performance in heterogeneously dense breasts (0.699-0.729). DenseNet121 achieved the highest overall performance (micro-averaged AUROC 0.885), and performance across categories was comparable between internal and external testing. For risk modeling, age combined with AI-derived density yielded a lower AUROC than age combined with mammography-reported density (0.541 vs. 0.570; p = 0.23), with no statistically significant difference. These findings indicate that deep learning models generalize well to external data with different racial composition for breast density assessment. While performance is strongest in extremely dense breasts, heterogeneously dense remains more challenging, highlighting the need for targeted optimization.

Computer Vision Scientific Discovery & Drug Design

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

External Validation of Deep Learning Models for BI-RADS Breast Density Prediction from Ultrasound Images

Related Papers