Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
1
Supervised fine-tuning outperforms complex reinforcement learning techniques in ensuring multilingual API reliability, challenging the notion that more sophisticated methods are always necessary.
English-only post-training of LLMs is suboptimal: even a single non-English language improves both English performance and cross-lingual generalization.