Search papers, labs, and topics across Lattice.
2
0
2
Achieving the first sub-Gaussian concentration bound for stochastic approximation under multiplicative noise could revolutionize how we approach stability in reinforcement learning algorithms.
Q-learning regret bounds can be achieved without optimism, but are highly sensitive to the suboptimality gap, motivating a new smoothed exploration strategy.