Search papers, labs, and topics across Lattice.
This paper introduces MCPEvol-Bench, a benchmark designed to evaluate the adaptability of LLM agents as they interact with dynamically evolving Model Context Protocol (MCP) servers. By employing 11 mutation operators across 123 MCP servers, the study reveals that even advanced models like GPT-5.4 and Claude-Sonnet-4-6 experience significant performance declines of 13.7% and 14.4%, respectively, when faced with evolving tool interfaces. These results underscore the critical need for robust assessments of LLM adaptability in changing environments, highlighting vulnerabilities in current LLM-driven workflows.
Even state-of-the-art LLMs show alarming performance drops when adapting to evolving toolsets, revealing a critical gap in current evaluation methods.
As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leverage real-world MCP servers to evaluate LLM agents'tool-using capabilities. However, these benchmarks overlook the continuous evolution of tool interfaces and functionalities within MCP servers, resulting in flawed assessments that fail to capture the agent's adaptability in changing tool landscapes. To bridge this gap, we introduce \textbf{MCPEvol-Bench}, a novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution. Inspired by large-scale empirical study, we propose 11 mutation operators to simulate realistic tool evolution within 123 MCP servers. We benchmark 12 state-of-the-art LLMs on multiple versions of MCP servers, revealing that even frontier models struggle to adapt to evolving tools. For instance, GPT-5.4 and Claude-Sonnet-4-6 exhibit performance declines of 13.7\% and 14.4\% in evolved MCP servers, respectively, accompanied by substantial increases in planning and reasoning errors. These findings highlight the vulnerability of LLM-driven workflows, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.