Search papers, labs, and topics across Lattice.
This study introduces OmnilingualGAIA2, a multilingual expansion of the GAIA2 agentic benchmark that evaluates AI agents across ten languages and five writing systems, addressing the critical issue of agentic competence transfer from English to other languages. The findings reveal a significant cross-lingual performance gap of 8.8-18.4 pass@3 points, primarily driven by model characteristics rather than translation issues, with a notable concentration of failures in tool orchestration tasks. The research highlights the necessity for multilingual evaluation in AI benchmarks to ensure equitable performance across diverse linguistic contexts.
A staggering 8.8-18.4 point performance gap in AI agents reveals that multilingual capabilities are not just a nice-to-have, but a critical oversight in current evaluations.
Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.