Search papers, labs, and topics across Lattice.
Shanghai AI Laboratory
2
0
3
GEPO achieves balanced cross-task improvements in RL for LLMs by dynamically adjusting advantages based on group entropy, outperforming traditional methods.
ThoughtFold cuts token usage by 56% without sacrificing accuracy by folding reasoning chains and eliminating redundant explorations.