Search papers, labs, and topics across Lattice.
Tencent
2
0
4
Self-improving search agents thrive when feedback and policy evolution are intertwined, leading to sustained performance gains and reduced hallucinations.
Stop wasting compute on full rollouts: ADWIN dynamically adapts on-policy distillation windows, slashing training costs by up to 4.1x without sacrificing accuracy on reasoning tasks.