HamburgUniversity of Southern DenmarkMay 21, 2026arXiv:2605.22734

ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning

Md Shamim Ahmed, Farzaneh Firoozbakht, Lukas Galke Poech, Jan Baumbach, Richard Röttger

AI Summary

The authors introduce ChronoMedKG, a temporally-grounded biomedical knowledge graph containing 460K+ evidence-linked triples covering 13K+ diseases, addressing the limitation of existing KGs in encoding temporal information crucial for clinical reasoning. They constructed the graph using a disease-autonomous multi-agent pipeline with multiple LLMs extracting knowledge from PubMed/PMC, ensuring multi-model consensus, credibility filtering, and ontology alignment. Experiments on the new ChronoTQA benchmark show that frontier LLMs struggle with temporal reasoning, and ChronoMedKG retrieval significantly improves their performance compared to HPOA-RAG.

Key Contribution

LLMs' clinical reasoning accuracy plummets by 30% when time matters, but a new temporally-aware knowledge graph recovers nearly two-thirds of that loss.

Abstract

Biomedical knowledge graphs (KGs) treat disease associations as static facts, but temporal information is crucial for clinical reasoning, e.g., a symptom diagnostic of one disease at age 3 may imply a different disease at age 13. Existing KGs such as PrimeKG, Hetionet, and iKraph do not encode when a finding becomes clinically relevant over the course of a disease. This limits their usefulness for longitudinal clinical reasoning and retrieval augmentation. We introduce ChronoMedKG, a temporal biomedical knowledge graph that contains 460,497 evidence-linked triples (filtered from 13M raw extractions) covering 13,431 diseases. Each association is tied to temporal components like onset window or progression stage, which are backed by PMID-traceable evidence and a multi-signal credibility score. The graph is constructed through a disease-autonomous multi-agent pipeline in which multiple frontier LLMs independently extract knowledge from PubMed and PMC literature. Only those relations are kept that are supported by multi-model consensus, survive credibility filtering, as well as ontology alignment. ChronoMedKG scored 92.7% agreement against Orphadata and adds temporal grounding for 6,250 diseases absent from HPOA, Orphadata, and Phenopackets, including 1,657 Orphanet-coded rare diseases. We further introduce ChronoTQA, a benchmark of 3,341 questions across eight task types (six temporal plus two static controls), with a 12-question supplementary probe. Frontier LLMs lose roughly 30 points moving from static to temporal questions; ChronoMedKG retrieval rescues 47-65% of their long-tail failures, against 17-29% for HPOA-RAG. As such, ChronoMedKG provides a crucial temporal axis for retrieval-augmented clinical systems that was previously absent.

Eval Frameworks & Benchmarks Recommendation & Information Retrieval Scientific Discovery & Drug Design

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning

Related Papers