Search papers, labs, and topics across Lattice.
Affiliation:, University of Maryland Baltimore County
4
1
8
5
Whether vision language models can decide when to answer immediately and, when more evidence is needed, which experiment to perform is studied to reveal limitations in physical reasoning that conventional answer accuracy can overlook.
OracleZoom is presented, an on-policy distillation-inspired, reference-constrained framework that trains on its trajectory while carrying the last ground-truth evidence beyond the supervision boundary, and achieves the state-of-the-art SR quality across zooming scales.
Machine translation benchmarks have functionally saturated, but pairing human-authored failure cases with deterministic verification rules reveals critical multimodal blind spots that automated metrics consistently miss.
VLMs are prone to critical failures that vary significantly across cultures, exposing the inadequacy of Western-centric safety benchmarks.