Search papers, labs, and topics across Lattice.
This paper introduces a meta-learning framework designed to enhance the alignment of large language models (LLMs) in multilingual contexts, particularly addressing the scarcity of human preference data in low-resource languages. By utilizing preference data from higher-resource languages, the framework facilitates a transferable initialization that allows for effective adaptation to target languages with minimal data input. The empirical results show significant improvements, with up to a 28% increase in win rates over baseline methods in extremely low-resource settings, confirming the robustness of the approach across various languages and model scales.
Achieving a 28% improvement in alignment performance with just 100 preference samples highlights the potential of meta-learning to bridge the data gap in multilingual LLMs.
Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language alignment, we propose a meta-learning framework for Reinforcement Learning from Human Feedback and Direct Preference Optimization. By leveraging preference data from other languages, our framework learns a transferable initialization that enables effective adaptation to a target language with minimal data. We provide theoretical guarantees for both the meta-reward modeling and meta-policy optimization settings, and empirically demonstrate the effectiveness of our approach on multilingual benchmarks. In an extremely low-resource setting with only 100 target-language preference samples, our approach achieves up to $28\%$ win-rate improvements over baseline methods, and consistently outperforms baselines across multiple target languages and model scales. Our approaches retain these advantages across different combinations of meta-training languages and varying linguistic distances from the target languages.