Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Mixed supervised fine-tuning outperforms next-chunk reasoning RL with significantly less compute, reshaping our understanding of effective training strategies for reasoning tasks.
PURA achieves over three times the message match rate of existing unbiased watermarking methods, all while preserving text quality and speed.