Search papers, labs, and topics across Lattice.
This study introduces R^3-Bench, a benchmarking framework that evaluates large language models (LLMs) on resource-rational reasoning tasks under shared budgets, contrasting with traditional independent budget assessments. The findings reveal that while an offline empirical oracle consistently outperforms the contest mean across various models and tasks, LLMs exhibit limited strategy updating and struggle to adapt under shared budget constraints. Notably, the results highlight a significant gap between the models' demonstrated single-problem competence and their performance in multi-problem scenarios, emphasizing the challenges in resource allocation strategies.
LLMs consistently underperform in shared-budget reasoning tasks, revealing a critical gap between their single-task capabilities and multi-task resource allocation.
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.