Search papers, labs, and topics across Lattice.
1
1
3
0
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.