Search papers, labs, and topics across Lattice.
3
1
6
9
RIPO redefines policy optimization for LLMs by correcting a fundamental geometric flaw, leading to unprecedented performance improvements in exploration efficiency.
Weak models can teach strong ones how to act better, boosting performance without the heavy lifting of direct RL training.
Extracting reasoning-effective updates from model geometry can dramatically enhance multi-domain performance while maintaining high reasoning capabilities.