Search papers, labs, and topics across Lattice.
To investigate cross-lingual consistency in culturally grounded alignment, the authors introduce C-Voices, a benchmark of 86,400 multilingual dilemma instances spanning 12 Chinese Social Value dimensions across six languages. Evaluating models on this benchmark reveals that value alignment diverges significantly across languages for identical dilemmas, exposing a fundamental failure mode in cross-lingual cultural coherence. To address this, they implement an inference-time activation steering approach that extracts value vectors from contrastive hidden-state discrepancies and intervenes on value-sensitive layers, enabling robust cross-lingual value steering without fine-tuning.
Identical ethical dilemmas trigger contradictory value judgments across languages in frontier LLMs, but inference-time steering vectors extracted from hidden-state discrepancies can reliably enforce culture-specific alignment without fine-tuning.
As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains underexplored, particularly for culturally grounded values, which are more abstract and difficult to evaluate and align than safety-centric principles. We investigate this issue through Chinese Social Values (CSV), a value system rooted in Chinese culture and comprising $12$ dimensions across national, societal, and personal levels. We construct C-Voices, the first comprehensive multilingual contrastive probe dataset for CSV, with 86,400 dilemma-based instances in six languages, each pairing a CSV-aligned action with a value-conflicting alternative. Building on the contrastive probes of C-Voices, we then propose a fine-tuning-free value vector steering method that derives value directions from hidden-state discrepancies and selectively intervenes on value-sensitive layers during inference. Experiments on six languages show that CSV-oriented preferences are model-dependent and language-sensitive, with the same dilemma eliciting divergent responses across languages. Our method achieves effective CSV steering, supports cross-lingual transfer of value vectors, and generalizes to existing FLAMES and ValuePrism.