Search papers, labs, and topics across Lattice.
This paper introduces the PL-NBA dataset, the first possession-level basketball video dataset designed to enhance visual understanding in sports by preserving the temporal continuity of game events. Comprising 11,000 offensive possession clips from 60 NBA games, it includes detailed annotations for 31,567 events, enabling complex tasks such as action anticipation and event recognition. Experimental evaluations reveal that current methods struggle with these tasks, highlighting PL-NBA as a significant benchmark for advancing sports video analysis.
Existing methods falter on a new benchmark that captures the full complexity of basketball game events, revealing critical gaps in current visual understanding approaches.
Visual understanding in sports has emerged as a hot topic in computer vision in recent years. Most existing basketball video datasets adopt single action or activity as sample, which can neither preserve the temporal continuity of game events nor support complex tasks such as action anticipation. To address this issue, this paper constructs the first possession-level basketball video dataset (PL-NBA), in which each sample is composed of a complete NBA offensive possession. Collected from 60 NBA games, PL-NBA contains 11,000 valid offensive possession clips and 31,567 annotated events with player names, captions, event types and timestamps. Each video clip includes multiple events and preserves the continuity of events, which is helpful for analysis of tactic. Experiment is conducted on multiple visual understanding tasks, including event recognition, video captioning, temporal action localization and action anticipation. Experimental results show that existing methods achieve limited performance on above four tasks, demonstrating that PL-NBA is a challenging benchmark for sports video understanding.