Search papers, labs, and topics across Lattice.
This paper introduces SynPre-FL, a novel framework that integrates high-fidelity synthetic electronic health record (EHR) generation with federated learning (FL) to enhance clinical risk prediction under non-IID conditions. By employing a latent autoencoder-diffusion model for synthetic data generation, the framework addresses challenges such as data scarcity, client heterogeneity, and class imbalance, while ensuring privacy protection against membership-inference attacks. Experimental results demonstrate that SynPre-FL significantly improves robustness and scalability in federated settings, yielding reliable and interpretable risk estimates across various client configurations.
Synthetic data can enhance federated learning for clinical risk prediction, achieving robust performance even in highly fragmented data environments.
Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks. Synthetic data generation may alleviate data scarcity, yet its integration with federated optimisation has received limited systematic study. We propose SynPre-FL, a unified framework combining high-fidelity synthetic EHR generation with synthetic-pretrained FL for robust prediction under non-IID conditions. A latent autoencoder-diffusion model generates privacy-preserving synthetic cohorts, which are used to warm-start federated training. This pretraining is followed by heterogeneity-aware optimisation using class-balanced local objectives, proximal regularisation, and adaptive server aggregation. Post-hoc calibration and federated-safe explainability support reliable and interpretable risk estimates. Experiments show that the synthetic generator preserves univariate, bivariate, and multivariate structure while protecting against membership-inference and reconstruction attacks. The generated data achieve strong downstream utility under TSTR, TRTS, and model-based evaluations. Across federated settings with 5, 10, and 15 heterogeneous clients, SynPre-FL consistently improves robustness and scalability over baseline methods, especially under severe non-IID fragmentation. Calibration improves probability reliability, while SHAP analysis produces stable and clinically coherent feature attributions across federation sizes. SynPre-FL therefore provides a practical and reproducible framework for combining synthetic data with FL to enable privacy-aware, interpretable, and robust clinical prediction from distributed tabular EHR data.