Search papers, labs, and topics across Lattice.
Affiliation:
2
1
4
0
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
OPD's "free lunch" of dense token-level reward may be an illusion, as teacher novelty, not just higher scores, drives successful distillation.