Search papers, labs, and topics across Lattice.
This paper introduces SpatialTrust, a benchmark designed to evaluate the capabilities of multimodal large language models (MLLMs) in recognizing and explaining environmental risks during secure authentication. The study reveals that existing MLLMs struggle significantly with indirect risk identification and explanation, underscoring a critical gap in their spatial risk awareness. Additionally, the authors present SpatialTrustGuard, a QA-and-audit pipeline that enhances the performance of the Qwen3-VL-30B-A3B-Instruct model, demonstrating the potential for structured methods to improve MLLM reliability in security contexts.
Current MLLMs can only identify and explain environmental risks in secure authentication scenarios with limited effectiveness, revealing a significant vulnerability in their design.
Visual environmental risk recognition plays an important role in secure authentication, where a user's surroundings may reveal sensitive information or introduce potential security risks. However, existing evaluations of multimodal large language models (MLLMs) rarely examine whether models can reliably recognize, localize, and explain such risks in spatially grounded authentication scenarios. We present SpatialTrust, a question-answering benchmark for evaluating environmental risk recognition in secure authentication. SpatialTrust assesses five complementary abilities: sensitive factor detection, direct factor identification, indirect factor identification, direct factor explanation, and indirect factor explanation. We evaluate both proprietary and open-source MLLMs and find that current models show limited performance, especially in understanding and explaining indirect risks, indicating that spatial risk awareness remains a challenging capability for MLLMs. In addition, we introduce SpatialTrustGuard, a structured QA-and-audit pipeline that improves Qwen3-VL-30B-A3B-Instruct from 36.78% to 41.12% overall. Our findings highlight the need for dedicated benchmarks and structured inference methods to improve the trustworthiness of MLLMs in secure authentication.