Search papers, labs, and topics across Lattice.
5
0
5
Bridging relevance and diversity in choice modeling, the DMNL bandit model achieves a new level of efficiency in assortment optimization.
Instance-optimality in MNL-based RL is now achievable, thanks to a new algorithm that adapts to the variance of learner-environment interactions.
Achieve statistically efficient regret guarantees in nonstationary bandit problems with *constant* per-round computation and memory costs – a feat previously unattainable.
MNL bandit optimization, previously plagued by computational bottlenecks, now has a scalable solution with provable optimality guarantees.
Overcome the generalization failures plaguing offline goal-conditioned RL, especially in long-horizon tasks, with a new latent representation alignment method that achieves SOTA results.