Search papers, labs, and topics across Lattice.
This paper introduces a reference-free framework utilizing large language model (LLM) judges to evaluate the quality of benchmarks for task-oriented conversational agents, focusing on consistency, complexity, and policy coverage. The framework is validated against human annotations and effectively distinguishes between benchmarks of varying quality, including those generated by LLMs and those subjected to quality-degrading perturbations. The findings highlight the critical need for robust evaluation methods in benchmark design, ensuring that assessments of conversational agents are reliable and meaningful.
LLM judges can reveal hidden flaws in conversational agent benchmarks, ensuring evaluations are both reliable and insightful.
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by LLMs of varying capabilities, as well as benchmarks subjected to controlled quality-degrading perturbations. Across domains and judge models, the proposed metrics consistently distinguish between benchmark quality levels. We further demonstrate the framework's applicability to manually curated benchmarks. Our framework offers a practical approach for evaluating synthetic and manually curated conversational-agent benchmarks.