Search papers, labs, and topics across Lattice.
York University
1
0
3
LVLM judges, despite excelling in English, exhibit surprisingly inconsistent and unreliable behavior when evaluating content in other languages, revealing a critical blind spot in current alignment and evaluation pipelines.