Search papers, labs, and topics across Lattice.
This paper introduces a computational argumentation-based framework for evaluating the faithfulness of LLM-generated summaries of parliamentary debates, focusing on preserving the argumentative structure related to policy proposals. The framework extracts and compares argument graphs from both the original debates and their summaries, assessing the preservation of key argumentative relationships. Results from a case study on European Parliament debates demonstrate the framework's ability to identify instances where LLM summaries fail to faithfully represent the original argumentative content.
LLM summaries of parliamentary debates often misrepresent the original arguments, but a new argumentation-based framework can pinpoint exactly where the reasoning goes wrong.
Understanding how policy is debated and justified in parliament is a fundamental aspect of the democratic process. However, the volume and complexity of such debates mean that outside audiences struggle to engage. Meanwhile, Large Language Models (LLMs) have been shown to enable automated summarisation at scale. While summaries of debates can make parliamentary procedures more accessible, evaluating whether these summaries faithfully communicate argumentative content remains challenging. Existing automated summarisation metrics have been shown to correlate poorly with human judgements of consistency (i.e., faithfulness or alignment between summary and source). In this work, we propose a formal framework for evaluating parliamentary debate summaries that grounds argument structures in the contested proposals up for debate. Our novel approach, driven by computational argumentation, focuses the evaluation on formal properties concerning the faithful preservation of the reasoning presented to justify or oppose policy outcomes. We demonstrate our methods using a case-study of debates from the European Parliament and associated LLM-driven summaries.