Search papers, labs, and topics across Lattice.
This study operationalizes Kahneman's dual-process theory to create Kahneman4Review, a benchmark for evaluating peer reviews based on nine textual dimensions and eight bias diagnostics across 3,563 rated reviews. The findings reveal a disconnect between LLM judges and human assessments of review quality, with LLMs favoring agentic reviews influenced by length and venue rather than genuine analytical quality. Additionally, the analysis indicates a temporal shift in review diagnostics coinciding with the rise of LLMs, highlighting the need for a more nuanced understanding of epistemic reliability in LLM-generated evaluations.
LLM judges may misinterpret peer review quality, favoring superficial traits over genuine analytical depth, raising questions about their reliability in academic assessments.
When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the difference matters philosophically. We operationalise Kahneman's dual-process theory into a structured rubric for peer review and release Kahneman4Review, a benchmark of 3,563 rated reviews scored along nine theoretically motivated textual dimensions, eight bias diagnostics, and a continuous reasoning-quality score. Three findings bear on trustworthiness: decision tier is not detectably aligned with the rubric's text-grounded epistemic-quality proxy; public-showcase agentic reviews receive higher raw scores than pooled human reviews, but length and venue explain most of the gap and the samples are not paper-paired; and ICLR review-text diagnostics shift at the 2022--2023 transition, temporally coincident with widespread LLM availability but without identifying its cause. A matched function-probe pilot further shows that the rubric distinguishes textual probes designed to contrast genuine fault-finding with surface fluency. We argue that a trustworthy reliability benchmark for LLM judges must separate analytical form from epistemic function, and propose concrete design choices toward that goal. An interactive demo is available at https://huggingface.co/spaces/nuojohnchen/Kahneman4Review.