Search papers, labs, and topics across Lattice.
This paper investigates statistical inference methods for quantile temporal difference learning (QTD) within the framework of distributional reinforcement learning, leveraging a generative model. The authors establish functional central limit theorems for both synchronous and asynchronous QTD, demonstrating that averaged iterates converge to a rescaled Brownian motion. They introduce an online inference method that utilizes random scaling to create an asymptotically pivotal statistic, significantly reducing memory requirements while enabling efficient inference throughout the QTD process.
Online inference in QTD can now be performed efficiently without the need to store entire trajectories, revolutionizing memory management in distributional reinforcement learning.
In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scaling, the inference procedure constructs an asymptotically pivotal statistic for inference by using the information along the whole QTD path. Meanwhile, the proposed statistic can be computed online without storing the entire trajectory of QTD iterates. This substantially reduces the memory requirement and enables efficient statistical inference in distributional reinforcement learning.