Search papers, labs, and topics across Lattice.
Hunyuan Team, Tencent
1
0
2
Static sample selection in reinforcement fine-tuning is a recipe for suboptimal updates鈥擠IEM adapts dynamically, leading to superior performance on reasoning tasks.