Search papers, labs, and topics across Lattice.
This paper introduces EgoArgus, a novel human-annotated dataset designed to benchmark vision-language models (VLMs) as situational assistants in real-world scenarios. The study reveals that current VLMs struggle to effectively arbitrate between visual and linguistic inputs, highlighting significant challenges in their reliability as egocentric assistants. Additionally, the findings indicate that existing methods for mitigating modality bias are limited, offering critical insights for practitioners aiming to deploy VLMs in daily assistance roles.
Current VLMs falter in their ability to discern trustworthy information from conflicting visual and linguistic inputs, raising questions about their reliability as daily assistants.
VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between visual evidence and user-provided language when the two are helpful, irrelevant, or conflicting. We introduce EgoArgus, a human-annotated dataset for evaluating egocentric assistants on understanding and decision tasks in five dialogue-video daily scenarios. Our results demonstrate that it is still challenging for current VLMs as reliable egocentric assistants, which requires identifying which modality is trustworthy and deciding when intervention is warranted. Deeper analysis also shows that existing modality bias mitigation methods are quite restricted to enhance performance, providing insights to aid practioners into the deployment of current VLMs as daily assistants.