Search papers, labs, and topics across Lattice.
To eliminate correlated errors where coding agents author flawed tests that validate their own broken patches, ExecCritic decouples test generation from code repair through an isolated test鈥搗erify鈥搑evise scaffold. Using Qwen-3.5-35B-A3B, the framework applies role-specific reinforcement learning to train a Test agent to generate discriminative, repository-native tests and a Repair agent to iterate against frozen, fail-closed execution feedback. When composed, the post-trained agents drive SWE-bench Verified resolution rates from an unguided 61.2% baseline to 72.6% without requiring oracle test feedback or proprietary frontier models at inference.
Naive agent-generated test feedback degrades SWE-bench performance by reinforcing shared hallucinations, but decoupling test generation from repair through role-specific RL converts a 3.9-point loss into an 11.4-point gain on open-weights models.
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair: a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from their execution feedback without changing the tests. Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In Learn to Test, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In Test to Improve, the Repair agent learns both direct task resolution and feedback-guided revision. On SWE-bench Verified, test quality determines whether feedback helps: holding the base Repair agent fixed, tests from the base Test agent reduce resolved rate from a no-test baseline of 61.2% to 57.3%, whereas tests from GPT-5.6-sol raise it to 65.3%. Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%; composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline without stronger-model or Oracle feedback at evaluation time. Code is publicly available at https://github.com/MSR-Orchard/execcritic.