Search papers, labs, and topics across Lattice.
2
0
3
A staggering 8.8-18.4 point performance gap in AI agents reveals that multilingual capabilities are not just a nice-to-have, but a critical oversight in current evaluations.
Standard MCQA metrics can mislead evaluations by over 2 points due to phrasing sensitivity, but ParaEval cuts this gap to under 1 point, revealing true model capabilities.