Search papers, labs, and topics across Lattice.
Affiliation:
5
0
3
Multilingual medical benchmarks penalize cross-lingual answer variation as model error, but real-world clinicians are sharply split on whether AI should enforce universal consistency or adapt to local cultural contexts.
MLLM-generated text shows distinct signs of translationese, but surprisingly, it diverges from traditional translation patterns in significant ways.
LLMs' internal confidence signals may mislead, as our verbalized methods reveal a stark disconnect between certainty and correctness in machine translation.
Instruction-tuned LLMs can be up to 26% overconfident in their own responses, but a simple adjustment can significantly improve their calibration.
LLMs fail spectacularly at understanding and generating Meenzerisch, a German dialect, achieving less than 10% accuracy even with targeted interventions, revealing a significant gap in their linguistic capabilities for low-resource languages.