Search papers, labs, and topics across Lattice.
This paper introduces SPRINT, a benchmark designed to evaluate proactive risk inference in multi-modal large language models (MLLMs) using a dataset of 2,888 sports videos, which includes both accident and safe control scenarios. The study reveals a significant disparity in MLLM performance, with the best model achieving over 95% hazard sensitivity but failing to identify causes accurately, with performance dropping below 50%. These results highlight the superficial nature of current MLLM safety capabilities and the urgent need for models that can provide reliable, cause-grounded early warnings in dynamic environments.
MLLMs may signal hazards with over 95% accuracy, yet they struggle to identify the underlying causes, revealing a critical gap in proactive safety capabilities.
Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a well-suited testbed: accident causes span diverse injury dimensions and pre-accident spatiotemporal cues draw on reasoning capabilities shared with broader safety domains such as autonomous driving and fall detection. We introduce SPRINT (Sports Proactive Risk INference Testbed), a benchmark of 2,888 real-world sports videos (2,440 accident, 448 safe controls) spanning 14 sports and 3 environmental settings. Accident videos feature fine-grained annotations of early hazard cues, accident timing, and hierarchical causes; safe videos are manually verified as accident-free and serve to diagnose prompt-induced false alarms. Evaluating state-of-the-art MLLMs under diverse prompts and temporal windows reveals a sharp gap between hazard sensitivity and understanding: the best model exceeds 95% in signaling hazards yet falls below 50% in identifying their causes. Diagnostic experiments further show that explicit danger queries trigger severe false alarms even on hazard-free videos. These findings indicate that current MLLMs exhibit only superficial proactive safety, lacking stable, cause-grounded early warning, and underscore the need for reliable proactive safety in dynamic physical environments. Data and code will be open-sourced upon acceptance.