Search papers, labs, and topics across Lattice.
The authors investigate how to account for valid lexical variation in document-level machine translation (MT) evaluation, introducing Cross-Term Variation (CTV) to measure whether variation relationships are faithfully preserved across languages. Evaluating four MT systems on English-French scientific corpora reveals that while glossary constraints improve standard accuracy and consistency scores, they degrade CTV by aggressively suppressing natural target-side variation. These findings show that standard consistency metrics conflate legitimate human-like variation with error, necessitating variation-aware metrics conditioned on source-side dynamics.
Strict glossary constraints artificially inflate translation consistency scores while actively destroying the natural lexical variation that human translators preserve.
Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source term. However, human translators routinely introduce variation that current metrics penalize as inconsistency. We examine how to account for this variation in document-level MT evaluation of English-French scientific translation, combining glossary-based accuracy, translation consistency, and a new cross-term variation (CTV) diagnostic measure that tests whether variation relationships are preserved across languages. Based on analyses of two parallel corpora, translated by four MT systems, we find that (1) MT systems generate less target-side variation than human translators; (2) transfer patterns strongly depend on the variation type; (3) consistency rankings vary with the choice of metric; and (4) constraining MT with a glossary improves accuracy and consistency but degrades CTV by suppressing valid variation. We argue for variation-aware evaluation that conditions consistency penalties on whether target-side variation mirrors source-side variation.