Search papers, labs, and topics across Lattice.
2
0
3
3
Execution time can be transformed into a learnable reward, leading to substantial improvements in code optimization performance in RL settings.
DecompRL enables LLMs to solve complex problems by breaking them down into manageable sub-tasks, achieving a 50x reduction in GPU costs while enhancing solution diversity.