Search papers, labs, and topics across Lattice.
This study investigates the performance of reasoning-oriented large language models (LLMs) on Theory of Mind (ToM) tasks, utilizing adaptations of machine psychological experiments alongside established benchmarks. The findings reveal that these reasoning models demonstrate enhanced robustness to variations in prompts and task perturbations, suggesting that their success in ToM tasks is more attributable to their general robustness rather than a unique ToM capability. This challenges the prevailing notion that LLMs possess a distinct understanding of ToM, highlighting the importance of robustness in their reasoning processes.
Reasoning-oriented LLMs may excel in Theory of Mind tasks not due to specialized abilities, but because they are more robust to variations in prompts and tasks.
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards have demonstrated notable improvements across a range of benchmarks. In this work, we examine the behavior of such reasoning models in ToM tasks using novel adaptations of machine psychological experiments together with results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis suggests these gains come at least partly from models being more robust at reaching the correct answer under prompt and task variation. We read this as evidence for a robustness-based account rather than for a new ToM-specific ability.