Search papers, labs, and topics across Lattice.
This study introduces OmniCSEval, a comprehensive benchmark designed to evaluate conversation summarization capabilities of LLMs across 1,800 diverse conversations and six real-world scenarios, addressing previous limitations in evaluation scenarios and sample sizes. Utilizing a bidirectional fact-checking framework, the authors assess LLMs on completeness, conciseness, and faithfulness, while also employing a human-LLM collaborative pipeline for enhanced accuracy. The findings highlight significant challenges faced by LLMs in cross-scenario performance, revealing insights into the interplay between reasoning capabilities, model scale, and efficiency in real-world applications.
Current LLMs struggle with cross-scenario summarization, revealing critical gaps in their reasoning capabilities and adaptability.
Despite the significant advancement of LLMs in conversation summarization, their evaluation remains limited by insufficient scenarios, input lengths, and sample sizes. Furthermore, existing benchmarks often omit frontier reasoning systems and efficient small models, or lack fine-grained, multi-dimensional assessments. To bridge these gaps, we propose OmniCSEval, a unified benchmark comprising 1,800 diverse conversations across six real-world scenarios, featuring context lengths ranging from 128 to 32k tokens. For fine-grained evaluation, we employ a bidirectional fact-checking framework that integrates key fact matching to assess completeness and conciseness, alongside summary fact verification to evaluate faithfulness. To ensure reliable assessment, we establish a human-LLM collaborative pipeline for key fact extraction and a multi-LLM consensus verifier for summary fact decomposition. Leveraging this framework, we evaluate 28 LLMs across four distinct categories grouped by reasoning capability and model scale. Our extensive empirical study reveals critical insights regarding the cross-scenario challenges current LLMs continue to face, the impacts of reasoning and scale, and the efficiency and adaptability of reasoning models. We also provide guidance for system selection in real-world deployments.