Search papers, labs, and topics across Lattice.
Affiliation:
4
0
4
5
High reasoning accuracy in VLMs doesn't equate to reliability, as shown by GPT-5.2's 96% hallucination rate despite top performance metrics.
LLMs can now be automatically audited for cultural insensitivity with 80% accuracy, thanks to a new metric that pinpoints and explains cultural errors in generated text.
Despite strong comprehension, LLMs still struggle to achieve human-level creativity in literary translation, often producing literal or contextually inappropriate renderings, especially when translating between distant languages.
MLLMs struggle to ground cultural values in visual scenes, losing ~7% accuracy compared to text-only inputs, even when they understand the visual content.