Search papers, labs, and topics across Lattice.
4
1
7
9
A new framework, Open-MOPD, boosts capability integration in multi-teacher distillation from 35.6% to 83.4% by addressing critical optimization imbalances.
RIPO transforms LLM reinforcement learning by correcting geometric flaws, leading to up to 60% better performance on key benchmarks.
Weak models can teach strong ones how to act better, boosting performance without the heavy lifting of direct RL training.
Extracting reasoning-effective updates from model geometry can dramatically enhance multi-domain performance while maintaining high reasoning capabilities.