Search papers, labs, and topics across Lattice.
This paper introduces a novel process-level benchmark for evaluating the agentic mathematical reasoning capabilities of Large Language Models (LLMs), addressing the limitations of traditional outcome-oriented assessments. By aligning problem-solving behaviors with a structured taxonomy of mathematical capabilities, the authors create a comprehensive suite of tasks that spans both textual and multimodal contexts. Experimental results indicate that LLMs with comparable end-to-end accuracy can have significantly different profiles in agentic capabilities, underscoring the importance of process-level evaluation in advancing LLM development.
Models that appear equally accurate can possess vastly different agentic reasoning capabilities, revealing the hidden complexities of LLM performance.
Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation provides limited diagnostic value for identifying process-level failures or rigorous logic, failing to guide the transformation of LLMs into robust agents. To bridge this gap, we present a process-level benchmark designed to evaluate the inherent agentic mathematical reasoning abilities of LLMs. Our framework aligns problem-solving agentic behaviors with a structured taxonomy of reusable mathematical atomic capabilities. We design a comprehensive suite of planning, action, and feedback tasks across both textual and multimodal contexts, supported by an automated pipeline that synthesizes high-quality trajectories and produces fine-grained annotations via controlled LLM rewriting. Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles. This demonstrates that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.