Search papers, labs, and topics across Lattice.
This paper introduces CORE-Bench, a novel benchmark designed to evaluate code retrieval in the context of agentic coding, which requires more than simple snippet matching. It addresses the limitations of existing benchmarks by incorporating three levels of evaluation: code understanding, issue-to-edit localization, and broader context retrieval, utilizing a dataset of over 180K queries. Experiments reveal a significant performance drop when transitioning from traditional code search to agentic coding settings, but also demonstrate that fine-tuning existing models can substantially enhance retrieval capabilities.
Traditional code retrieval methods falter in agentic coding, but CORE-Bench reveals that fine-tuning can bridge the gap and boost performance significantly.
Code retrieval is becoming central to coding agents, but agentic coding requires more than matching a natural-language query to an isolated snippet. Given a user request, a coding agent needs to navigate a concrete repository state, locate relevant files and functions, gather supporting context, and filter similar in-repository distractors. Existing code retrieval benchmarks mainly evaluate docstring-to-function or snippet-level matching, thereby missing this requirement-driven repository search problem. To address this gap, we introduce CORE-Bench, a comprehensive benchmark for code retrieval in the era of agentic coding. CORE-Bench evaluates code retrieval ability at three levels: code understanding, issue-to-edit localization, and broader context retrieval. Built from curated code-search tasks and SWE-bench-series instances, CORE-Bench contains over 180K queries and 106K broader-context relevance labels. Experiments with representative embedding models show a sharp drop from traditional code search to code retrieval in agentic coding settings. Simple supervised fine-tuning of existing embedding models significantly improves performance in this setting, suggesting substantial room for further progress.