Search papers, labs, and topics across Lattice.
Meituan LongCat Interaction, Peking University
2
0
4
OCSD reveals how isolating observation effects can lead to more effective token-level updates in reinforcement learning, outperforming traditional methods.
LLM reasoning gets a serious upgrade with MASPO, a new RLVR method that smartly balances gradient use, probability mass, and signal reliability for faster, more robust learning.