Search papers, labs, and topics across Lattice.
This study introduces SportD, a benchmark designed to evaluate the strategic decision-making capabilities of vision-language models (VLMs) in soccer by analyzing 478 on-ball decisions from the 2022 FIFA World Cup. The models' choices are compared against a possession-value model that quantifies the optimal action for increasing scoring probabilities, revealing that the best-performing VLM selects the highest-valued action only 31.4% of the time, significantly lower than professional players at 38.9%. The findings indicate that VLMs tend to favor lower-risk, lower-reward actions, suggesting a reliance on familiar patterns rather than optimal strategic evaluation.
VLMs struggle with strategic decision-making in soccer, showing a preference for safer plays over optimal actions, which could reshape our understanding of their reasoning capabilities.
Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use information to make strategically effective decisions. We investigate this question in soccer, where models observe the seconds preceding an on-ball decision and must choose whether to shoot or pass to a specific teammate. Unlike conventional visual-understanding tasks, soccer enables decisions to be evaluated quantitatively by estimating the value of every available action. We introduce SportD, a benchmark comprising 478 on-ball decisions from the 2022 FIFA World Cup. Each model choice is evaluated against a possession-value model that estimates the action that most increases the attacking team's probability of scoring, allowing us to measure both optimal-action accuracy and the value forfeited by suboptimal decisions. Across three frontier VLMs, the best selects the highest-valued action on 31.4% of events, compared with 38.9% for the professional players, and all models incur significantly greater regret. Further analysis reveals a systematic preference for lower-variance and lower-reward actions: VLMs shoot less often and select substantially less progressive passes than either the optimal policy or the real players. The models also reproduce the player's specific action above chance even when that action is suboptimal, suggesting partial imitation of familiar play patterns rather than consistent evaluation of counterfactual alternatives. SportD provides a value-grounded testbed for measuring physical strategic reasoning in VLMs.