Search papers, labs, and topics across Lattice.
2
0
2
Horizon-free regret can be achieved with a new algorithm that dramatically outperforms previous methods by eliminating logarithmic horizon dependence.
Forget RL fine-tuning: this paper shows you can beat it at cold-start personalization with a tiny model and clever Bayesian inference over structured preference priors.