Search papers, labs, and topics across Lattice.
3
0
5
4
Zero-shot VLMs falter at geometric reasoning, with only one out of five models surpassing random performance on basic jigsaw puzzles, revealing a scaling cliff in their capabilities.
Reasoning-aware retrieval can boost language model performance by surfacing diverse solution strategies that traditional methods overlook.
Current visual grounding models struggle to infer objects from contextual roles and intentions, highlighting a critical gap in their ability to perform true scene understanding.