Search papers, labs, and topics across Lattice.
1
0
1
Rigorous offline confidence intervals for optimal policy values are finally achievable over diverging horizons and non-stationary data, eliminating the classic non-smoothness bottleneck of maximum Bellman operators.