Search papers, labs, and topics across Lattice.
This study introduces FairGlucose, a comprehensive benchmark for evaluating the fairness of continuous glucose monitoring (CGM) AI tools across diverse patient demographics, utilizing a cohort of 300 patients with 132,480 forecasting samples. The analysis of 33 models reveals that while aggregate performance metrics appear stable, significant disparities exist at the subgroup level, particularly with type 1 diabetes (T1D) patients experiencing higher prediction errors compared to type 2 diabetes (T2D) patients. These findings indicate that population-level validation can mask critical inequities, underscoring the necessity for subgroup-disaggregated reporting in digital health AI assessments.
Population-level validation of CGM AI tools hides critical disparities, with T1D patients facing 6 mg/dL higher prediction errors than T2D counterparts.
As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p<0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.