Search papers, labs, and topics across Lattice.
Affiliation:
4
0
8
Code generation models can now self-improve at test time without ground-truth unit tests by turning behavioral execution agreement on generated input probes into a stable, hack-resistant policy gradient signal.
TCS not only generates sound test cases but also adapts to the model's weaknesses, leading to a marked improvement in code generation performance.
Achieving high-quality model performance with just 10% of the required labels could revolutionize the scalability of RLVR in large language models.