Search papers, labs, and topics across Lattice.
12
2
11
12
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
OPDVR transforms the landscape of model distillation by ensuring that only correct trajectories enhance learning, leading to significant performance gains on reasoning tasks.
Only 5.4% of coding agent attempts successfully complete a migration while preserving behavioral correctness, exposing a critical gap in current AI capabilities for software evolution.
Agents struggle to significantly improve training algorithms, with the best only achieving 25% of the potential optimization gap.
A new framework, Open-MOPD, boosts capability integration in multi-teacher distillation from 35.6% to 83.4% by addressing critical optimization imbalances.
Self-evolving agents are failing to adapt effectively in dynamic environments, with top methods achieving less than 70% success on benchmark tasks.
Weak models can teach strong ones how to act better, boosting performance without the heavy lifting of direct RL training.
Browser agents can achieve unprecedented scalability by harnessing the collective skills of internet users through skill distillation.
AI coding agents excel at translating scientific tasks into familiar formats but struggle to achieve true scientific discovery, with only 17.8% surpassing state-of-the-art benchmarks.
Sustained self-improvement in LLM agents is achievable through a novel adaptive framework that outperforms traditional methods in dynamic task environments.
Intrinsic reward signals in unsupervised RL for LLMs inevitably collapse due to sharpening of the model's prior, but external rewards grounded in computational asymmetries offer a path to sustained scaling.