Search papers, labs, and topics across Lattice.
This study evaluates the effectiveness of multi-agent debate tools in providing feedback on research papers compared to a single-pass report from a frontier AI model. In a controlled experiment involving authors of 44 meta-analyses, the single-pass report was preferred by authors over the multi-agent debate outputs, despite the latter consuming significantly more resources. The findings suggest that multi-agent debate may not enhance AI feedback quality, raising concerns about the reliability of AI judges in academic contexts.
Authors favored a single AI report over complex multi-agent debates, challenging the assumption that more voices lead to better feedback.
Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors'place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.