Search papers, labs, and topics across Lattice.
2
0
3
Naive agent-generated test feedback degrades SWE-bench performance by reinforcing shared hallucinations, but decoupling test generation from repair through role-specific RL converts a 3.9-point loss into an 11.4-point gain on open-weights models.
TRACE transforms how long-horizon agents are trained, leading to a remarkable performance increase on complex tasks without the need for supervised fine-tuning.