Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
Standard PPO clipping actively starves the rare, correct reasoning chains models need most during rollout reuse, a failure mode resolved by replacing hard clipping thresholds with asymmetric, $\alpha$-divergence-derived sequence weighting.