Search papers, labs, and topics across Lattice.
This paper develops a novel theory of debate judgement that addresses how outcomes of agentic AI debates are evaluated, focusing on properties such as reproducibility, robustness, groundedness, and explainability. The authors compare two debate judgement methods鈥擫LMs as judges and formal semantics from computational argumentation鈥攄emonstrating that while both achieve similar accuracy, the latter provides stronger formal guarantees. This work highlights the potential of argumentation semantics as a foundational framework for improving debate-driven AI systems.
Argumentation semantics could be the key to reliable and explainable debate judgement in AI, outperforming traditional LLM-based methods in formal guarantees.
Debates have recently emerged as a useful methodology for agentic AI to improve performance as well as to aid explainability and user engagement. For example, LLM-empowered agents may debate internally (with themselves) and/or externally (with other agents). In many settings where debates are used, debates' outcomes and resulting outputs are determined post-hoc by external judges, often LLMs. In this paper we develop and test a novel theory of debate judgement applicable to all settings where agents engage in debates by providing pros and cons for their opinions therein. Specifically, we identify a number of formal properties that debate judgement may be required to satisfy in general, as concerns reproducibility, robustness, groundedness and explainability. Then, we explore their satisfaction formally and/or experimentally, for claim verification settings, for two specific alternative debate judgement methods: variants of the LLMs as a judge idea and formal semantics drawn from computational argumentation. We show that the two methods give similar accuracy performances but the former may lack formal guarantees that the latter brings. Overall, our study indicates argumentation semantics as an ideal candidate for principled judges in debate-driven AI.