Search papers, labs, and topics across Lattice.
Baidu Inc, University of Science and Technology of China
3
0
6
ECHO enables RL agents to retain and leverage fine-grained historical evidence, achieving a 43.4% accuracy on complex tasks while using less context than prior methods.
Stop uniformly distilling your LLMs: SCOPE selectively amplifies teacher guidance on incorrect trajectories and reinforces student uncertainty on correct ones, leading to significant gains in reasoning performance.
LLM reasoning gets a serious upgrade with MASPO, a new RLVR method that smartly balances gradient use, probability mass, and signal reliability for faster, more robust learning.