Search papers, labs, and topics across Lattice.
This paper introduces a novel characterization of the occupancy measure in reinforcement learning by embedding the planning criterion into the dynamics through a resetting planning process, resulting in a new stationary measure called the visitation measure. The authors reveal that the achievable visitation measures form a dually flat statistical manifold, allowing for a generalized approach to planning-as-inference that extends from linear rewards to nonlinear functionals. Key findings include the interpretation of temporal-difference error as a marginal-utility estimate, which enhances the understanding of decision-making in reinforcement learning and theoretical neuroscience.
Achievable visitation measures in reinforcement learning form a dually flat statistical manifold, transforming our understanding of planning-as-inference.
We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.