Search papers, labs, and topics across Lattice.
6
1
11
10
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
Frontis-MA1 achieves a remarkable 71.21% Medal Average on MLE-Bench Lite, outperforming leading models and showcasing the potential of AI systems to recursively improve their own engineering processes.
OPD's "free lunch" of dense token-level reward may be an illusion, as teacher novelty, not just higher scores, drives successful distillation.
LLMs can achieve massive performance gains on reasoning and knowledge-intensive tasks simply by iteratively refining their answers using pseudo-labels derived from unlabeled data.
Intrinsic reward signals in unsupervised RL for LLMs inevitably collapse due to sharpening of the model's prior, but external rewards grounded in computational asymmetries offer a path to sustained scaling.