Search papers, labs, and topics across Lattice.
2
0
4
28
Standard GRPO's uniform clipping boundary actively chokes exploration by penalizing rare breakthroughs on hard problems just as harshly as trivial rollouts on easy ones.
Reverse-KL objectives can provably avoid catastrophic forgetting in continual learning, unlike the commonly used forward-KL.