Search papers, labs, and topics across Lattice.
This study explores the application of Vision Language Models (VLMs) for annotating video game frame sequences with reward signals, addressing the challenges posed by the variability of synthetic scenarios and their alignment with real-world physics. The authors find that VLMs frequently fail to accurately respond to fundamental queries in racing games, indicating limitations in their current capabilities across various genres. By analyzing factors such as input sequence length, resolution, and question batching, the researchers propose strategies like output mixing and prompt optimization to enhance annotation quality.
VLMs struggle to provide accurate annotations in video games, revealing significant gaps in their understanding of dynamic environments.
Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization. We also show how input sequence length, resolution, and question batching affect the annotation quality and its token consumption.