Search papers, labs, and topics across Lattice.
This paper introduces the Knowledge Synthesis Review (KSR) framework, which benchmarks large language models (LLMs) like GPT-5 and Claude Sonnet 4 on their ability to perform multi-source evidence synthesis tasks such as screening, extraction, analysis, and synthesis. The evaluation, conducted on a diverse 244-document subset, revealed that while no single model excelled across all tasks, Claude Sonnet 4 achieved the highest screening accuracy at 82.8% and GPT-5 demonstrated superior recall at 91.8%. The findings highlight the necessity of human oversight, particularly in interpretive analysis and cross-source synthesis, underscoring the framework's potential to uncover critical insights that single-source approaches might overlook.
No LLM can dominate all tasks in evidence synthesis, revealing the critical need for a human-in-the-loop approach to ensure comprehensive analysis.
Evidence in rapidly evolving fields is fragmented across academic studies, industry reports, policy documents, and media sources that differ in quality, structure, and purpose, making timely synthesis difficult. Large language models (LLMs) may accelerate this work, but their reliability across the distinct cognitive tasks of a review remains uncertain. We introduce the Knowledge Synthesis Review (KSR), a human-in-the-loop framework that decomposes evidence synthesis into screening, extraction, analysis, and synthesis, benchmarks LLM-based systems on each task against expert reference standards, and routes each task to the best-performing system under continuous expert validation. We evaluated GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, and NotebookLM on a 244-document benchmark subset drawn from a 1,893-document corpus on AI and work spanning four source types, against a gold standard with high inter-rater reliability (92.2% agreement, kappa = 0.80). No system led on all tasks. Claude Sonnet 4 achieved the highest screening accuracy (82.8%) and GPT-5 the highest recall (91.8%) at the expense of lower specificity. Extraction exceeded 90% agreement for titles and sources but degraded in author and reference fields. Performance declined most in interpretive analysis and cross-source synthesis, where expert judgment remained essential. A contamination check on post-cutoff documents showed no evidence that prior exposure inflated results. Applied to the full corpus, the routed workflow surfaced cross-source asymmetries and blind spots that single-source synthesis would miss, including worker well-being, small firms, and the Global South. KSR offers a transparent, auditable, model-agnostic framework for governing LLM assistance in research synthesis while preserving human accountability.