Search papers, labs, and topics across Lattice.
Meituan LongCat Interaction, Fudan University
2
0
4
OCSD reveals how isolating observation effects can lead to more effective token-level updates in reinforcement learning, outperforming traditional methods.
Teacher guidance can be strategically enhanced by targeting high-disagreement states, leading to significant performance gains in agentic tasks.