Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
3
Across four models, attention heads causally necessary for OCR are identified, and it is discovered that these are in fact general-purpose heads that output interpretable semantic features across all image tokens.
Knowledge retrieval and reasoning are the main bottlenecks in KI-VQA, but a new diagnostic benchmark reveals deeper issues in visual grounding and object identification.