Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Multilingual medical benchmarks penalize cross-lingual answer variation as model error, but real-world clinicians are sharply split on whether AI should enforce universal consistency or adapt to local cultural contexts.
Instruction-tuned LLMs can be up to 26% overconfident in their own responses, but a simple adjustment can significantly improve their calibration.