Search papers, labs, and topics across Lattice.
Microsoft Research
1
0
2
Standard next-token prediction on an agent's self-generated explanations outperforms GRPO on SWE-bench in half the training updates鈥攅ven bootstrapping on tasks where every initial rollout failed.