Search papers, labs, and topics across Lattice.
This paper elucidates the foundational properties of the Bellman equation by establishing that its recursive structure stems from three interrelated conditions: sufficient statistics in dynamics, recursive return decomposition, and compatible uncertainty aggregation. The authors demonstrate that when these conditions are satisfied, the Bellman equation emerges from their mutual consistency, while violations can often be mitigated through state augmentation or modifications to return or dynamics. Additionally, the framework reveals three dualities鈥攂etween probability and return, return and aggregation, and aggregation and probability鈥攖hereby unifying disparate methodologies across reinforcement learning, control, and decision theory.
The Bellman equation's structure isn't just a mathematical artifact; it emerges from deep interdependencies that can reshape our approach to decision-making in AI.
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one between probability and return, one between return and aggregation, and one between aggregation and probability. Our framework reveals these dualities as arising from a single construction, unifying methods developed separately across reinforcement learning, control, and decision theory.