Search papers, labs, and topics across Lattice.
This study introduces the "render ceiling," a model-free benchmark for evaluating vision-language models by rendering known objects and recovering answers through cross-view correspondence. The authors demonstrate that this method effectively isolates misreading from misreasoning, revealing that deficiencies in model performance are inherent rather than artifacts of evaluation. Their findings show that while providing exact geometry as text improves performance across fourteen models, a purely supervised vision model outperforms all vision-language models, indicating a significant gap in current multimodal approaches.
A model-free benchmark reveals that vision-language models often misread rather than misreason, exposing a performance gap that a supervised vision model can surpass.
Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separating the two places a second model in the loop. We introduce the render ceiling, a model-free reference for benchmarks built by rendering known objects: inverting the frozen cameras and re-solving cross-view correspondence recovers exactly the answer the images support. We prove the ceiling fails only through an enumerable set of projection coincidences and certify that set empty on 2,160 rendered crystal structures, so every point of a model's deficit belongs to the model. Across fourteen vision-language models, supplying exact geometry as text lifts every model yet closes under half the gap for thirteen, while a supervised vision model with no language component reads the same images at 0.8952, above every vision-language model. The instrument exposes extraction-stage fabrication that downstream accuracy would misattribute to reasoning, yields camera-placement rules for benchmark builders, and transfers to any benchmark with an invertible forward rendering.