Search papers, labs, and topics across Lattice.
1
0
3
15
Naive RL fine-tuning for code generation can lead to LLMs regurgitating the same solutions, but penalizing code similarity boosts performance even more than directly optimizing for pass@k.