Search papers, labs, and topics across Lattice.
York University
2
0
3
LVLM judges, despite excelling in English, exhibit surprisingly inconsistent and unreliable behavior when evaluating content in other languages, revealing a critical blind spot in current alignment and evaluation pipelines.
Finally, a reliable, reference-free metric exists to evaluate factual consistency in Bangla summarization, unlocking progress in this under-resourced language.