Search papers, labs, and topics across Lattice.
2
0
4
STARE not only prevents policy entropy collapse but also enhances accuracy by up to 8% across diverse tasks, showcasing a new frontier in stable RL training.
LLMs can reason more effectively by directly tracking their own belief in the correct answer throughout the reasoning process, enabling more targeted policy updates.