Search papers, labs, and topics across Lattice.
4
7
3
15
Achieving a 28% improvement in alignment performance with just 100 preference samples highlights the potential of meta-learning to bridge the data gap in multilingual LLMs.
Even with corrupted human feedback, surprisingly tight guarantees for multi-agent reinforcement learning are possible.
RLHF and DPO are surprisingly vulnerable to data poisoning, with even a small number of carefully crafted preferences capable of steering the learned policy towards a desired (potentially harmful) target.
RLHF models can be made significantly more robust to distribution shift by incorporating distributionally robust optimization into both reward modeling and policy optimization.