Search papers, labs, and topics across Lattice.
Meta AI
1
0
2
Off-Context GRPO achieves a 3.9% absolute improvement in reasoning tasks by leveraging privileged information to guide learning, even when models struggle to produce correct outputs.