Search papers, labs, and topics across Lattice.
This study introduces XIH-Bench, a novel benchmark designed to evaluate instruction hierarchy (IH) compliance in multilingual large language models (LLMs) across six languages and various domains. The findings reveal a language-dependent asymmetry in IH compliance, where certain languages enhance compliance when prioritized but hinder it when demoted, alongside a surprising Language Boundary Effect that shows cross-language conflicts yield higher compliance than same-language ones. These results highlight significant multilingual reliability and security risks, especially due to language specialization affecting the override of lower-priority instructions.
Language-dependent asymmetries in instruction hierarchy compliance reveal that multilingual models can be more reliable in cross-language contexts, but risk security vulnerabilities due to language specialization.
Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones. Despite its importance for safe and controllable deployment, existing evaluations have focused almost exclusively on English, leaving it unclear whether IH compliance remains stable in multilingual settings. We introduce XIH-Bench, a benchmark for multilingual IH evaluation with both same-language and cross-language conflicts across six languages, four domains, and three IH settings. Across models, we find two consistent patterns. First, IH compliance exhibits a clear language-dependent asymmetry: a language that strengthens compliance in the higher-priority position can become disruptive in the lower-priority position. Second, cross-language conflicts yield higher compliance than same-language conflicts, a phenomenon we term the Language Boundary Effect. We further show that language specialization can make lower-priority instructions in model-favored languages harder to override, creating multilingual reliability and security risks.