Search papers, labs, and topics across Lattice.
Wangxuan Institute of Computer Technology, Peking University
2
0
4
Achieving a 35% reduction in depth RMSE, NSL-SLAM not only enhances tracking accuracy but also ensures robust performance in real-world scenarios without catastrophic failures.
RGPO doesn't just reweight samples like PPO; it *rejects* the bad ones, leading to a Pareto-dominant improvement in reward and KL divergence during RLHF.