Search papers, labs, and topics across Lattice.
1
0
Achieving up to $\tilde{\mathcal{O}}(T_{total}^{-1})$ convergence rates without regularization challenges conventional wisdom about the necessity of entropy in policy gradient methods.