Search papers, labs, and topics across Lattice.
This study evaluates three open-source vision-language models (VLMs) for their ability to classify egocentric robot images into four levels of proxemic danger, which is essential for safe navigation in human environments. The researchers compared various prompting strategies and fine-tuning methods, finding that while fine-tuning yielded only modest improvements overall, the model \textit{Qwen-VL} with an advanced prompt significantly outperformed others in recalling high-danger cases. Importantly, the analysis revealed that accurate danger classification does not necessarily correlate with effective spatial grounding, highlighting limitations in current VLMs' proxemic reasoning capabilities.
\textit{Qwen-VL} can achieve significantly better recall for high-danger situations, but its success doesn't guarantee accurate spatial awareness.
Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (\textit{InternVL}, \textit{Qwen-VL}, and \textit{SmolVLM}) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tuning against a stratified random baseline. Without fine-tuning, all models perform near the baseline, while fine-tuning yields only modest overall improvements. However, \textit{Qwen-VL} with an advanced prompt achieves substantially higher recall for high-danger cases than the other models. An analysis of person localization further shows that correct danger classification does not correspond to better spatial grounding, indicating that a model may produce a useful safety label without attending to the relevant region of the scene. These results show that current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, although targeted prompting and fine-tuning can improve high-danger detection in selected models.