Search papers, labs, and topics across Lattice.
2
0
5
4
Training agents in deep, evolving environments can dramatically enhance their performance, with a 9B model achieving a 30.6% accuracy increase through targeted design.
Achieve up to 12x greater sample efficiency in reasoning tasks by relaxing strict imitation constraints in on-policy distillation, enabling smaller models to match the performance of much larger ones.