CUHKResearcherSJTUUIUCXJUJun 5, 2026arXiv:2606.07297

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Shaoqiu Zhang, Yuhang Wang, Jialiang Liang, Yuling Shi, Wenhao Zeng, Maoquan Wang, Shilin He, Ningyuan Xu, Siyu Ye, Kai Cai, Xiaodong Gu

AI Summary

This paper introduces SWE-Explore, a novel benchmark designed to evaluate the repository exploration capabilities of coding agents, addressing the limitations of existing benchmarks that overlook fine-grained agent capabilities. By analyzing 848 issues across 10 programming languages and 203 open-source repositories, SWE-Explore provides a structured evaluation of how well coding agents can retrieve relevant code regions under a fixed line budget. The findings reveal that agentic explorers significantly outperform classical retrieval methods, particularly in line-level coverage and efficient ranking, which are critical for effective bug diagnosis and code localization.

Key Contribution

Coding agents can achieve superior repository exploration, outperforming classical methods by effectively leveraging line-level context for bug diagnosis and code retrieval.

Abstract

Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tasks as a holistic, binary prediction problem (e.g., resolved or unresolved), neglecting fine-grained agent capabilities such as repository understanding, context retrieval, code localization, and bug diagnosis. In this paper, we introduce SWE-Explore, a benchmark that isolates the evaluation of repository exploration, a critical capability of coding agents. Given a repository and an issue, SWE-Explore asks an explorer to return a ranked list of relevant code regions under a fixed line budget. SWE-Explore covers 848 issues across 10 programming languages and 203 open-source repositories. For each instance, we derive line-level ground truth from independent agent trajectories that successfully solved the same issue, distilling the specific code regions their solution paths actually consulted. We evaluate exploration along coverage, ranking, and context-efficiency dimensions, showing that these metrics strongly track downstream repair behavior. Across a broad set of retrieval methods, general coding agents, and specialized localizers, we find that agentic explorers form a clear tier above classical retrieval. While file-level localization is already strong for modern methods, line-level coverage and efficient ranking remain the key axes differentiating state-of-the-art explorers.

Code Generation & Program Synthesis Eval Frameworks & Benchmarks Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Related Papers