Search papers, labs, and topics across Lattice.
This paper introduces ReViCo, a benchmark designed to assess the capabilities of Vision Language Models (VLMs) in understanding and correcting text within images. By evaluating models on their ability to identify and rectify text errors in real-world visual contexts, the study reveals significant performance gaps between VLMs and human capabilities, indicating that current models struggle with visual text perception. The findings underscore the need for improved methodologies in VLM development to enhance text-aware functionalities.
VLMs exhibit a significant performance gap in visual text understanding, with even the best models falling short of human accuracy in error correction tasks.
Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo (Real Visual Correction), a benchmark designed to evaluate VLM text understanding through a novel task of visual text error correction. ReViCo challenges models to identify and fix text errors in real-world images, which requires a profound understanding of the interplay between visual text and its surrounding visual context. We benchmark various VLMs using two distinct paradigms: prompt-based strategy and targeted model training, both aimed at pushing the limits of current models. Our experiments reveal a striking performance gap between even the best VLMs and human, and further analysis also shows that most models struggle to accurately perceive the visual text, resulting in frequent correction errors. By highlighting these gaps, ReViCo provides a new benchmark foundation for developing more robust and text-aware VLMs.