Search papers, labs, and topics across Lattice.
1
2
4
Mixed SFT outperforms next-chunk reasoning RL in reasoning tasks while using over 60 times less compute, reshaping our understanding of effective training strategies.