Search papers, labs, and topics across Lattice.
4
0
4
4
Models may score well on benchmarks but often fail to meet strict perceptual requirements, revealing a hidden brittleness in multimodal evaluations.
Multimodal LLMs still struggle to faithfully recreate webpages from videos, particularly in capturing fine-grained style and motion, despite advances in other areas.
VLMs can't count blocks because they lack a view-consistent spatial interface, but decomposing scenes into orthographic projections fixes it.
Even reward models that get the right answer can be dangerously wrong in their reasoning, leading to worse RLHF outcomes, but R-Align fixes this by explicitly aligning rationales with gold standard judgments.