Search papers, labs, and topics across Lattice.
This paper establishes a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) within the framework of tabular distributional reinforcement learning. By employing a global comparison argument that leverages the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, the authors demonstrate how to bring an arbitrarily initialized iterate into a local neighborhood, where a variance-sensitive martingale analysis can be applied. The key result indicates that the leading last-iterate fluctuation is of order $\widetilde O\bigl(T^{-a/2}/\sqrt{1-\gamma}\bigr)$, revealing a clear distinction between local stochastic fluctuations and global sample complexity.
The leading last-iterate fluctuation in quantile temporal-difference learning can be controlled without polynomial dependence on the number of quantiles, challenging conventional assumptions about sample complexity.
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes $\alpha_t=c(t+1)^{-a}$ with $a\in(1/2,1)$, the leading last-iterate fluctuation is of order $\widetilde O\bigl(T^{-a/2}/\sqrt{1-\gamma}\bigr)$ and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order $m^{-1}$ in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.