Search papers, labs, and topics across Lattice.
University of Alberta
2
0
5
LVLM judges, despite excelling in English, exhibit surprisingly inconsistent and unreliable behavior when evaluating content in other languages, revealing a critical blind spot in current alignment and evaluation pipelines.
LLMs exhibit a measurable preference for American English due to skewed pretraining data, inefficient tokenization of British English, and biased generative outputs.