Search papers, labs, and topics across Lattice.
This paper introduces EgoSafe-Bench, a novel benchmark designed to evaluate visual safety understanding in first-person scenarios by focusing on causal reasoning under epistemic uncertainty. The benchmark consists of 12,000 samples generated through a Hierarchical Reasoning Evaluation (HRE) protocol, which enforces logical consistency and penalizes superficial reasoning. Evaluations of state-of-the-art Large Vision-Language Models (LVLMs) reveal a significant gap between high descriptive performance and poor causal reasoning, highlighting the need for improved logical robustness in video understanding systems.
LVLMs may excel at describing scenes but often falter in causal reasoning, revealing a critical gap in their understanding of first-person visual safety.
Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate impressive semantic alignment on standard benchmarks, they often struggle to distinguish between superficial correlation and genuine forensic logic when grounded in the dynamic, partially observable nature of first-person experiences. Existing evaluations, dominated by third-person surveillance footage and binary classification metrics, fail to expose this cognitive gap. To address this, we introduce EgoSafe-Bench, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios. It comprises 12,000 unique evaluation samples, generated by pairing each of the 3,000 video clips with a QA chain governed by our proposed Hierarchical Reasoning Evaluation (HRE) protocol. Unlike standard benchmarks, HRE mandates a rigorous reasoning trajectory from initial feature anchoring to blind-spot deduction and intent inference, thereby enforcing logical consistency and penalizing shortcut-based predictions.Extensive evaluations of state-of-the-art LVLMs (e.g., Qwen3-VL, Gemini, VideoLLaMA 3) reveal a significant perception-reasoning decoupling: models often achieve high descriptive scores but exhibit notable fragility in causal reasoning and logical closure. Our work provides both a challenging dataset and a systematic evaluation framework to foster the development of logically robust video understanding systems.