Search papers, labs, and topics across Lattice.
1
0
2
6
RGPO doesn't just reweight samples like PPO; it *rejects* the bad ones, leading to a Pareto-dominant improvement in reward and KL divergence during RLHF.