Search papers, labs, and topics across Lattice.
This study introduces PatchBench, a novel benchmark designed to rigorously evaluate AI agents' performance in vulnerability patching by addressing two critical threats: patch memorization and superficial fixes. By implementing a patch similarity metric, the research reveals that approximately 25% of agent-generated patches closely resemble historical developer patches, highlighting significant concerns regarding the validity of current evaluation methods. The findings indicate that existing validation techniques inflate the perceived effectiveness of AI agents, with a 1.83脳 increase in solve rates when relying solely on Proof-of-Concept tests, underscoring the need for more robust evaluation frameworks in vulnerability repair.
One in four AI-generated patches may simply replicate historical fixes, raising serious questions about the reliability of current vulnerability patching evaluations.
AI agents have recently demonstrated strong performance in automated vulnerability patching. However, existing evaluations often validate a patch only by testing whether the provided Proof-of-Concept (PoC) input still triggers a crash. This leaves two key threats to validity: agents may reproduce memorized historical developer patches, or they may generate surface-level fixes that only suppress the reported crash. We study these concerns for C/C++ vulnerability patching. We introduce a patch similarity metric to detect memorized patches. On average, 25% of the agent patches exhibit substantial similarity to historical developer patches, indicating that patch memorization is a real threat to the validity of vulnerability patching evaluations. Meanwhile, agents also frequently exploit benchmark structures to pass patch validation by patching on the crash stack trace to suppress the crash, rather than localizing and fixing the root cause of the vulnerabilities. To handle these issues, we propose PatchBench, a new benchmark for evaluating AI agents on realistic vulnerability patching tasks. PatchBench selects vulnerabilities whose ground-truth fixes lie outside the crash stack and uses vulnerability transplant and code mutations to migrate historical vulnerabilities into new repository contexts, reducing the risks of surface-level fixes and patch memorization. We develop new patch validation methods that thoroughly evaluate both security and semantic correctness of agent patches. Across 11 state-of-the-art agents, including the top three AIxCC agents, the original PoC-only validation inflates the patching task solve rate of agents by 1.83$\times$ on average. Our results reveal key limitations of current patching agents and point to future research directions for more reliable vulnerability repair.