Search papers, labs, and topics across Lattice.
This paper introduces evidence-ledger adjudication, a novel workflow that enhances claim-evidence traceability by pairing claims with evidence packets and assessing their support relations. Using a comprehensive benchmark of 2,335 rows derived from multiple datasets, the method significantly outperforms traditional baselines, achieving 0.676 relation accuracy and 0.601 macro-F1 scores. The approach effectively routes unsupported or contradictory claims back to authors, demonstrating its potential to create an auditable layer for AI-assisted writing.
Evidence-ledger adjudication enables AI to not only draft claims but also ensure their accuracy by effectively tracing and validating the supporting evidence.
AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-agent baseline. It also routes 1270/1435 claims whose gold labels indicate contradiction, missing evidence, or mixed evidence, while routing 295/900 supported claims. These results show that evidence-ledger adjudication can turn heterogeneous evidence packets into an auditable traceability layer for AI-assisted writing.