Search papers, labs, and topics across Lattice.
5
1
8
12
Weak models can teach strong ones how to act better, boosting performance without the heavy lifting of direct RL training.
Browser agents can achieve unprecedented scalability by harnessing the collective skills of internet users through skill distillation.
AI coding agents excel at translating scientific tasks into familiar formats but struggle to achieve true scientific discovery, with only 17.8% surpassing state-of-the-art benchmarks.
Sustained self-improvement in LLM agents is achievable through a novel adaptive framework that outperforms traditional methods in dynamic task environments.
Intrinsic reward signals in unsupervised RL for LLMs inevitably collapse due to sharpening of the model's prior, but external rewards grounded in computational asymmetries offer a path to sustained scaling.