Search papers, labs, and topics across Lattice.
3
3
3
62
Optimistic Multi-step Preference Optimization is built upon the optimistic online mirror descent algorithm and provides a rigorous analysis for the convergence of OMPO and shows that OMPO requires O ( ϵ − 1 ) policy updates to converge to an ϵ -approximate Nash equilibrium.
Raven achieves superior long-context recall by intelligently routing memory updates, outperforming traditional models that struggle with interference and eviction.
Forget state-action spaces: this work achieves efficient multi-agent imitation learning by concentrating on feature-level representations in linear Markov games.