Search papers, labs, and topics across Lattice.
The authors tackle offline statistical inference for the optimal policy value by replacing the non-differentiable maximum Bellman operator with a smooth softmax counterpart and deriving two novel nuisance functions via self-induced Bellman equations. Valid confidence intervals for optimal values are notoriously difficult to obtain offline due to non-smoothness and distributional shift across long horizons. By constructing a Neyman-orthogonal debiased estimator, they establish asymptotic normality under diverging horizons and time-varying behavior policies, requiring only standard statistical convergence rates for the underlying machine learning estimators.
Rigorous offline confidence intervals for optimal policy values are finally achievable over diverging horizons and non-stationary data, eliminating the classic non-smoothness bottleneck of maximum Bellman operators.
We study offline inference for the optimal value in reinforcement learning. Two new nuisances are derived as fixed points of a self-induced Bellman equation, in which we approximate the maximum Bellman operator by its softmax correspondence. We propose a debiased estimator through the Neyman orthogonality and establish its asymptotic normality under diverging horizons even when the behavior policy changes with time, as long as the nuisances have the statistical rates that can be achieved by many machine learning methods. We provide a concrete estimating procedure for these nuisances and show they can lead to valid inference. Synthetic experiments validate the numerical performance of our inference method, and we implement it in real-life decision-making problems, including bike repositioning and AI agentic tool use.