Search papers, labs, and topics across Lattice.
This study conducts a comparative analysis of generative models for synthetic transcriptomic data, focusing on the integration of biological knowledge through gene graphs to enhance data quality. The research highlights the limitations of existing methods due to biases and constraints in real datasets and introduces MK-TGAN, a multi-kernel, Graph Neural Network-based model that excels in generating realistic and biologically plausible synthetic data. The findings demonstrate that incorporating prior knowledge significantly improves the performance of synthetic data generation, making it more applicable for downstream biomedical tasks.
MK-TGAN outperforms traditional generative models by integrating biological knowledge, producing synthetic transcriptomic data that is both realistic and useful for research.
As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data generation can help overcome these limitations. Here, we present a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs. This ensures that synthetic data capture real-world gene patterns, maintaining their usefulness for downstream tasks. In particular, we introduce and benchmark three variants of the Generative Adversarial Network. Among the alternatives, MK-TGAN - an innovative multi-kernel, Graph Neural Network-based model - stands out for its performance in terms of both the realism and utility of the generated data. Unlike other methods, MK-TGAN leverages prior knowledge graphs by exploiting graph neural networks. Our results show that prior knowledge integration strategies improve performance, and that MK-TGAN consistently produces synthetic samples with superior realism and biological plausibility.