Search papers, labs, and topics across Lattice.
This study develops and evaluates four automated segmentation pipelines for detecting lesions associated with age-related macular degeneration (AMD) and diabetic macular edema (DME) in optical coherence tomography (OCT) images, achieving Dice scores between 0.76 and 0.82. The research emphasizes the importance of generalization to clinical data by testing the models on an external cohort, OLIVES, using a novel proxy-metric framework that correlates predictions with clinical biomarkers. Key findings indicate that while the models can track lesion burden outside the training distribution, their performance is not as robust as in-domain evaluations, suggesting a need for further refinement in clinical applications.
Automated segmentation pipelines can effectively track AMD and DME lesions in clinical settings, but their performance drops when generalizing beyond training data.
Age-related macular degeneration (AMD) and diabetic macular edema (DME) are leading causes of vision loss, and optical coherence tomography (OCT) is the standard modality for detecting and monitoring the subtle lesions that drive treatment decisions. Most deep-learning segmentation work for OCT is validated only in-domain, leaving generalization to clinical data collected under different acquisition protocols largely untested. This work develops and systematically ablates four lesion-segmentation pipelines -- 2D and 3D variants for AMD and DME -- reaching Dice scores of 0.76 to 0.82 with strong volumetric and surface calibration (r vol, r surf greater than or equal to 0.97 across all four pipelines) on an in-domain validation set. The ablation process establishes a full-volume, calibration-aware adoption standard that catches mechanisms an ordinary slice-level evaluation would keep, and identifies ensemble composition as the most consistent driver of improvement. To test generalization, the models are evaluated on OLIVES, an external clinical cohort with no lesion-level ground truth, using a proxy-metric framework built around biomarker AUROC, central subfield thickness (CST) correlation, and longitudinal concordance. Predictions track clinical biomarkers outside the training distribution, though less strongly than in-domain -- evidence for, not validation of, automated lesion-burden tracking as a clinical tool.