Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Naive agent-generated test feedback degrades SWE-bench performance by reinforcing shared hallucinations, but decoupling test generation from repair through role-specific RL converts a 3.9-point loss into an 11.4-point gain on open-weights models.