Search papers, labs, and topics across Lattice.
2
0
4
Standard next-token prediction on an agent's self-generated explanations outperforms GRPO on SWE-bench in half the training updates鈥攅ven bootstrapping on tasks where every initial rollout failed.
Context compaction in agentic RL can run up to 5x faster simply by streaming the KV cache instead of flushing it鈥攁ccidentally turning standard LLMs into recurrent agents that preserve evicted context purely through RL.