Search papers, labs, and topics across Lattice.
This study evaluates the effectiveness of federated training for tokenized generative event models (GEMs) using a dataset of over 122,000 intensive care hospitalizations from three health systems. The results indicate that GEMs not only outperform conventional supervised models in terms of transportability across sites but also achieve competitive performance with centralized training, particularly in scenarios with limited local data. The findings highlight the potential of federated learning approaches to maintain high predictive accuracy while addressing the challenges posed by siloed electronic health records.
Federated training of generative event models can achieve near-centralized performance while significantly enhancing cross-site transportability in healthcare data.
Electronic health record foundation models are limited by institutionally siloed data and substantial performance degradation under cross-site transfer. We evaluated federated training of tokenized generative event models (GEMs) across 122,251 intensive care hospitalizations from three independent health systems harmonized to the Common Longitudinal ICU Data Format. Models were assessed on 12 post-24-hour clinical prediction tasks using within-site, cross-site, centralized, and federated training configurations. GEMs achieved the highest mean within-site and cross-site ROC-AUC and were substantially more transportable than conventional supervised models: their average cross-site penalties were 0.025 ROC-AUC and 0.027 PR-AUC, compared with 0.079 and 0.089 for LightGBM. Federated Learning (FedAvg and FedAvgM) approached the performance of centralized GEM training, with most gains obtained within 5-10 communication rounds. However, centralized multi-site training provided only modest improvements over complete local training. Multi-site models were most useful when local training data were limited, with their advantage narrowing as institutional data accumulated. These findings show that federated GEM training is technically feasible and preserves most centralized performance, but that the main open challenge is learning transportable representations to translate larger, but heterogeneous data from multiple health systems into a reliable target-site benefit.