Search papers, labs, and topics across Lattice.
This paper introduces BasketballBench, a multimodal benchmark designed to evaluate comprehensive basketball understanding through 7,980 questions across ten tasks involving text, image, and video data. The benchmark addresses the limitations of existing evaluations that assess abilities in isolation, revealing that current multimodal large language models (MLLMs) struggle with tasks requiring the integration of multiple capabilities. In contrast, the proposed BasketballSkills agent, which utilizes eight basketball-specific perception and retrieval tools, significantly outperforms MLLMs, demonstrating the value of structured, domain-specific skill composition for enhanced understanding of basketball dynamics.
Current MLLMs falter in integrating multiple basketball knowledge domains, while BasketballSkills showcases superior performance through structured skill composition.
Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA season and includes official playby-play, rosters and profiles for 530 active players, and 2,501 possession-level broadcast clips. We further propose BasketballSkills, an agent that composes eight basketball-specific perception and retrieval tools under four reusable skills that specify tool order, evidence bindings, and stopping conditions. Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.