Search papers, labs, and topics across Lattice.
This paper addresses the challenge of untranslatability in machine translation (MT) by introducing a structured ontology and a taxonomy of compensation strategies to convey meaning when direct translation is impossible. The authors operationalize this framework into a multilingual dataset of untranslatable sentences with corresponding strategy-based translations, facilitating controlled analysis of translation behavior. Initial studies reveal that translation quality significantly varies based on the chosen strategy, with a preference for translations that incorporate explanatory context, highlighting the importance of strategic approaches in MT systems.
Translation quality hinges on the strategy employed, with human preferences favoring context-rich explanations over direct equivalence.
Untranslatability, cases where meaning cannot be directly preserved across languages, is well-studied in linguistics but underexplored in NLP. As machine translation (MT) systems improve on standard benchmarks, their limitations increasingly concentrate in such cases, where translation cannot be reduced to one-to-one equivalence. We introduce a structured ontology of untranslatability along with a taxonomy of compensation strategies, which are specific techniques to convey meaning under these untranslatable circumstances. We operationalize this framework into a multilingual dataset of untranslatable sentences paired with strategy-based translations, enabling controlled analysis of translation behavior. Initial human preference studies suggest that translation quality depends on the strategy used, with consistent preferences for outputs that include explanatory context, known as the Annotation compensation strategy. Our framework and dataset provide a foundation for studying and modeling strategy-informed machine translation.