Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
SRPO enables LLMs to self-reflect and transform sparse feedback into dense learning signals, achieving state-of-the-art performance with drastically reduced training costs.