Search papers, labs, and topics across Lattice.
This study investigates the structural sensitivity of multilingual large language models (LLMs) to semantics-preserving perturbations, specifically focusing on Hindi and Malayalam. By employing two perturbation methods鈥攃onstrained constituent reordering and active-passive voice transformation鈥攐n a newly introduced benchmark dataset, IndicReStruct, the authors reveal that even state-of-the-art LLMs exhibit significant degradation in mathematical reasoning performance when faced with structurally altered inputs. The findings highlight critical weaknesses in entity-quantity alignment and suggest that intermediate transformer layers are pivotal for reasoning restoration, indicating a lack of robust compositional invariance in these models.
Multilingual LLMs suffer substantial reasoning failures when faced with structurally altered inputs, revealing a critical vulnerability in their design.
Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to semantics-preserving structural variation remains underexplored, particularly for relatively free word-order languages. We investigate the structural sensitivity of multilingual LLMs using two linguistically grounded perturbation settings in Hindi and Malayalam: constrained constituent reordering and active-passive voice transformation. We introduce a benchmark dataset IndicReStruct, with two variants, GSM8K-Reordered and GSM8K-Voice, constructed from GSM8K while preserving semantic meaning. Across six state-of-the-art LLMs and multiple prompting strategies, we observe consistent and significant degradation in mathematical reasoning performance under structurally perturbed inputs. To further understand these failures, we perform qualitative error analysis and mechanistic interpretability experiments using residual-stream activation patching. Our analyses show that reasoning failures frequently arise from disruptions in entity-quantity alignment and that intermediate transformer layers contribute most strongly toward reasoning restoration. Overall, our findings suggest that current multilingual LLMs remain highly sensitive to surface syntactic realization and lack robust compositional invariance under structurally different but semantically equivalent inputs.