Search papers, labs, and topics across Lattice.
This paper introduces Cultivar, a novel translation benchmark that emphasizes source-contrastive evaluation to address issues of contamination and localization in multilingual translation. By evaluating 32 open-weight models against both localized and unlocalized datasets, the authors reveal significant performance discrepancies that highlight the models' varying robustness across different locales. Key findings indicate that machine translation (MT) specialized models exhibit less robustness and tend to favor translations of US-centric content, raising concerns about their generalizability and cultural sensitivity.
MT models show a troubling bias towards US content, with specialized models proving less robust across diverse locales.
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.