Search papers, labs, and topics across Lattice.
Tencent
4
0
7
TRACE redefines rollout budget allocation by treating each turn in a multi-turn interaction as a unique node, leading to improved reward contrast and policy performance.
Stop wasting compute on full rollouts: ADWIN dynamically adapts on-policy distillation windows, slashing training costs by up to 4.1x without sacrificing accuracy on reasoning tasks.
Most RLVR datasets are just remixes of a few originals, and this paper shows how to trace them back to their source, revealing widespread data contamination.
Training tool-calling agents with just an 8B language model outperforms traditional methods that depend on expensive resources, reshaping the landscape of tool learning.