Search papers, labs, and topics across Lattice.
This paper critiques the conventional wisdom that more criticisms yield better reviews in AI-assisted evaluations, highlighting the risks of omission and overcritique. The authors introduce EquiReview-R, a novel approach that refines a structured set of concerns based on evidence, effectively distinguishing between different types of review failures. The results demonstrate a significant reduction in major overcritique from 15.5% to 8.1% while maintaining a low rate of omissions, showcasing the efficacy of their evidence-guided refinement process.
EquiReview-R reduces major overcritique by nearly half while ensuring critical issues are not overlooked, challenging the notion that quantity of criticism equates to quality.
AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures obscure the distinction. We therefore recast AI-assisted review as evidence-guided refinement of a structured concern set, with omission and overcritique treated as separate risks. Building on this formulation, we introduce EquiReview-R, which resolves existing concerns against localized evidence, searches for missing issues from independent and review-conditioned perspectives, and returns stop, continue, or defer. To expose the failure mode that motivates this design, we construct an evidence-linked trajectory corpus. Its retrospective analysis shows why revision must precede further search: nearly all concerns in a high-recall review lack a definitive evidential disposition, while an earlier refinement mechanism cannot revise them. On a frozen cohort of previously unseen papers, EquiReview-R satisfies the prespecified non-inferiority criterion for major omission, reduces major overcritique from 15.5% to 8.1%, and attains a one-sided omission upper bound of 9.9% while stopping on 52.4% of papers. Computation-matched controls, controlled pairs, and ablations show that the gain comes from revision rather than extra inference or shorter output. We release the corpus as ReviewTrace, an evidence-linked resource for studying review revision, disagreement, and provenance.