Search papers, labs, and topics across Lattice.
This paper introduces DoublesEval, a diagnostic framework designed to assess multi-agent tactical reasoning in Vision-Language Models (VLMs) using professional doubles badminton as a testbed. By employing a key-moment-based protocol, the framework evaluates models across four dimensions of reasoning, revealing significant weaknesses in spatial state understanding and interaction binding. The proposed TacticCheck method improves model performance by reranking answers based on lower-level tactical predictions, yet substantial gaps in robust reasoning persist across all tested models.
VLMs struggle with multi-agent tactical reasoning, revealing critical weaknesses in spatial understanding and interaction binding that current models fail to address.
Visual Language Models (VLMs) excel at describing visible scene content but struggle to reason about dynamic multi-agent interactions, where action semantics depend on coordinated roles and spatial-temporal dependencies. We formalize this capability as \textbf{multi-agent tactical reasoning} and introduce \textbf{DoublesEval}, a diagnostic evaluation framework that leverages professional doubles badminton as a structurally tractable testbed. DoublesEval employs a key-moment-based protocol that decomposes rallies into tactically salient instants and probes models across four interpretable dimensions: atomic recognition, intra-segment composite understanding, cross-segment causal reasoning, and high-level tactical abstraction. This design isolates \emph{where} reasoning fails, rather than merely measuring answer correctness. To address observed failure modes, we propose \textbf{TacticCheck}, a lightweight constraint-guided test-time consistency checker that reranks candidate answers using the model's own lower-level tactical predictions, requiring no parameter updates or ground-truth labels at inference time. Evaluating four representative open-source VLMs on 60 curated rallies (yielding $\sim$9.6K structured instances) via a zero-shot protocol, we find that models remain weak across all diagnostic levels, with especially clear bottlenecks in spatial state, interaction binding, and terminal evidence. TacticCheck delivers consistent gains across all evaluated models, while still leaving a substantial gap to robust tactical reasoning. These results highlight the need for structured, interaction-aware evaluation paradigms for next-generation VLMs. The source code is available in \href{https://github.com/Chengjt1999/DoublesEval}{\textcolor{blue}{our GitHub repository}}.