Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
0
Mixed-policy reinforcement learning can enable language models to absorb knowledge more effectively than traditional supervised fine-tuning, especially in complex reasoning scenarios.