Search papers, labs, and topics across Lattice.
This paper introduces PACE-Bench, a benchmark designed to evaluate the performance of self-evolving agents in dynamic environments by linking source and mutated target environments across six physics domains. The study finds that existing self-evolving methods struggle to adapt effectively, with the best-performing method achieving only 66.7% success on a subset of tasks, highlighting the limitations of current adaptation strategies. The results indicate that while simulator-grounded reflection is more effective than unverified self-revision, significant challenges remain, suggesting a need for redesigning mechanisms rather than merely tuning parameters.
Self-evolving agents are failing to adapt effectively in dynamic environments, with top methods achieving less than 70% success on benchmark tasks.
Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchmark of 144 source-to-target adaptation pairs across six physics domains. Each pair links a source environment to a mutated target environment with the same goal and interface. A code-driven design that succeeds in the source fails in the target, where agents must iteratively adapt it into a working target design using diagnostic sandbox feedback within a limited attempt budget. We compare ten self-evolving methods from four paradigms. The benchmark remains far from saturated: Reflexion + Qwen3-14B succeeds on only 35.9\% of full-benchmark pairs, while GPT-5.5 solves 66.7\% of the Statics subset under the full budget. Together, these results show that simulator-grounded reflection is more reliable than unverified self-revision, while memory anchors agents to early designs and broad tree search explores without converging. Even revealing exact physical changes does not raise the performance ceiling, pointing to mechanism redesign rather than parameter inference as the central bottleneck. Data and code are available at https://github.com/thunlp/PACE-Bench.