Search papers, labs, and topics across Lattice.
3
0
4
2
Execution time can be transformed into a learnable reward, leading to substantial improvements in code optimization performance in RL settings.
Extrapolating between code-generating RL agents trained on different unit test coverages unlocks better correctness-efficiency trade-offs than any single agent alone.
LLMs can learn to "debug" their own code by simulating execution, leading to significant gains in competitive programming performance.