Search papers, labs, and topics across Lattice.
This paper introduces MCR-Bench, a novel benchmark designed to evaluate the multi-round interactive nature of real-world code reviews, addressing the limitations of existing static approaches. By analyzing 2,269 real-world code review tasks across five programming languages, the authors reveal that mainstream large language models (LLMs) struggle with defect detection and lifecycle state tracking, particularly as interaction rounds increase. Key findings indicate that LLMs exhibit significant performance variability based on defect complexity and severity, highlighting critical weaknesses in their handling of temporal alignment and long-range memory.
Mainstream LLMs miss 40% of defects in multi-round code reviews, revealing their limitations in real-world software development contexts.
In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming. Although recent work explores large language models (LLMs) for automated code review, most approaches oversimplify code review into a single-round, static decision task, which fails to capture the multi-round interactive nature and the complex problem-solving processes inherent in realistic review scenarios. To bridge this gap, we introduce MCR-Bench, the first defect state-aware benchmark designed for realistic multi-round code review. MCR-Bench covers five commonly-used programming languages and consists of 2,269 real-world multi-round code review tasks, each of which is annotated with fine-grained defect information and cross-round state labels. Each task in MCR-Bench is equipped with fine-grained defect metadata (e.g., description, type, severity) alongside dynamic state annotations, capturing the complete evolutionary trajectory of a defect throughout the multi-round process. We obtain several findings through extensive experiments on MCR-Bench with mainstream LLMs. (1) Limited overall capability: experiments reveal that mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases; (2) Defect-sensitive performance: LLMs'performance varies substantially across different defect types and severity levels, with semantically complex or low-salience defects being significantly more likely to be missed; (3) Underlying Failure Mechanisms: our in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory.