Search papers, labs, and topics across Lattice.
Tencent AI Lab
3
0
8
21
The effectiveness of on-policy distillation hinges not on the size of the teacher model, but on the quality of the guiding signal, revealing critical pathologies that can undermine exploration.
Self-distillation in a verifiable environment enables web agents to achieve competitive performance without reliance on external teacher models.
Learnable critics that evaluate the model's own GUI grounding proposals, rather than relying on static geometric heuristics, unlock substantial gains in accuracy.