Search papers, labs, and topics across Lattice.
This paper redefines Key Point Analysis (KPA) as a structured prediction problem, emphasizing the need for semantic groupings and prevalence estimation in generating representative key points. The authors identify significant limitations in existing KPA benchmarks, including issues with grouping quality and coverage, which hinder effective evaluation. By introducing a structure-aware, distribution-sensitive benchmark through human-in-the-loop re-annotation, they demonstrate that their approach yields superior coherence and reliability in key point extraction compared to current methods.
Existing KPA benchmarks fail to deliver reliable evaluations, but a new structure-aware benchmark reveals significant improvements in coherence and quality of key points.
Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. We argue that KPA is fundamentally a structured prediction problem that requires recovering semantic groupings, generating representative key points, ensuring coverage, and estimating prevalence. Under this formulation, we show that existing KPA benchmarks suffer from limitations in grouping quality, redundancy, coverage, and argument-key point mappings, causing ceiling violation and selection failure in reference-based evaluation. To support future research on true KPA, we introduce a structure-aware, distribution-sensitive benchmark built via a human-in-the-loop re-annotation. Human and LLM evaluations consistently show that the resulting structures yield more coherent groupings, higher-quality key points, better coverage, and more reliable prevalence estimates than existing annotations. We further release several annotation resources to support research on KPA evaluation, argument-key point matching, explainable KPA, and LLM-as-a-judge methodologies, and outline a research agenda for true KPA.