Search papers, labs, and topics across Lattice.
This study introduces a framework that leverages biomedical knowledge graphs (bioKGs) to enhance the prioritization of candidate biological annotations for expert review. By employing relation-specific binary classifiers trained with a community-based negative sampling strategy, the authors achieve more reliable confidence estimates for annotation plausibility. Experimental results across five bioKGs reveal a 5.8% average increase in balanced accuracy, demonstrating that this approach not only improves classifier robustness but also facilitates more efficient AI-assisted biomedical curation while maintaining expert oversight.
Prioritizing candidate biomedical annotations with a novel framework using bioKGs boosts classifier accuracy and efficiency in expert curation.
The rapid growth of biomedical knowledge has made the validation of automatically generated biological annotations a major bottleneck in biomedical curation. While computational methods can rapidly produce large numbers of candidate annotations, determining which are biologically valid still requires costly expert review. Prioritizing these candidates before manual curation has therefore become a fundamental challenge. Machine learning techniques can support this process by exploiting biomedical knowledge graphs (bioKGs), which capture biological entities and their functional associations. In this work, we propose a framework that leverages bioKGs to estimate the plausibility of candidate annotations and guide expert curation. Starting from knowledge graph embeddings, we train relation-specific binary classifiers using a community-based negative sampling strategy to obtain reliable confidence estimates. We then introduce a family of plausibility measures that combine classifier confidence, classifier reliability, and the semantic context provided by alternative relationships involving the same pair of biological entities. Unlike conventional confidence estimation, the proposed approach explicitly accounts for multiple biologically meaningful relations that may coexist between the same entities. Experimental results on five large bioKGs demonstrate that the proposed negative sampling strategy consistently improves classifier robustness, increasing balanced accuracy by an average of 5.8%. Moreover, the plausibility measures outperform classifier confidence alone, enabling more effective prioritization of candidate annotations for expert review. Overall, our results show that the use of bioKGs improves the efficiency of AI-assisted biomedical curation while preserving expert control over the final annotation assessment.