Search papers, labs, and topics across Lattice.
Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, of Space Integrated Information System
2
0
4
GUPO reveals that accounting for gradient uncertainty can dramatically improve policy optimization in post-training LLMs, leading to more effective reasoning capabilities.
Temporal GRPO reveals that aligning reinforcement learning updates with task stages can significantly boost both efficiency and success rates in complex vision-language-action tasks.