Search papers, labs, and topics across Lattice.
This paper investigates the fixed-policy distributional soft Bellman operator within the framework of distributional soft policy iteration (DSPI) using Cramér geometry, a cumulative distribution function (CDF)-based metric. The authors demonstrate that this operator exhibits a contraction property, leading to a unique fixed point that facilitates convergent iterative policy evaluation. By establishing a connection between the CDF formulation and the spectral domain, the study provides valuable insights into the evaluation error and critic-loss design in DSPI algorithms.
The Cramér-geometric Bellman operator reveals a unique fixed point that could transform how we approach evaluation errors in distributional reinforcement learning.
Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns. Theoretical analysis of such an evaluation step requires a probability metric under which Bellman updates can be controlled, typically by showing that the operator contracts the distance between any two candidate return-distribution estimates. In this paper, we focus on the Cramér geometry, a cumulative distribution function (CDF)-based metric with an $L^2$ structure, and study whether the fixed-policy distributional soft Bellman operator has this contraction property and hence a unique fixed point under this metric. Working directly on an admissible CDF field domain, we formulate the CDF-level distributional soft Bellman operator, prove that it is a $\sqrtγ$-contraction, and obtain the corresponding unique fixed point together with convergent iterative policy evaluation. The CDF formulation also shows that this finite-Cramér-domain property follows from a uniform first-moment condition on the combined one-step reward entropy shift, rather than from separate uniform boundedness assumptions on the reward and entropy terms. We then transport the same evaluation problem to the spectral domain by conjugation, obtaining an equivalent Hilbert-space representation of the same decision process. Taken together, these results identify the Cramér-geometric Bellman fixed point associated with the policy-evaluation step of DSPI, providing a reference point for studying approximate critics, evaluation error, and critic-loss design in DSPI-style algorithms.