Search papers, labs, and topics across Lattice.
3
0
3
2
Models may score well on benchmarks but often fail to meet strict perceptual requirements, revealing a hidden brittleness in multimodal evaluations.
Multimodal LLMs still struggle to faithfully recreate webpages from videos, particularly in capturing fine-grained style and motion, despite advances in other areas.
Even reward models that get the right answer can be dangerously wrong in their reasoning, leading to worse RLHF outcomes, but R-Align fixes this by explicitly aligning rationales with gold standard judgments.