Search papers, labs, and topics across Lattice.
This paper introduces Visual Credit Audit (VCA), a method that evaluates the contribution of images in multimodal spatial reasoning tasks by distinguishing between the support provided by images versus text-only contexts. The study finds that a significant percentage of correct decisions (12.73-26.25%) are made without proper credit to the images, highlighting a potential oversight in current evaluation metrics. By employing relation-specific visual evidence and controlled experiments, VCA reveals that traditional benchmarks may misrepresent model capabilities, particularly in relation to the visual evidence utilized in decision-making.
A staggering 12.73-26.25% of correct decisions in multimodal spatial reasoning are made without proper credit to the supporting images, exposing flaws in current evaluation methods.
Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.