Search papers, labs, and topics across Lattice.
2
0
1
Bridging the gap between reinforcement learning and control theory could unlock new synergies in optimizing unknown dynamical systems.
Q-learning can provably converge to optimal policies in non-Markovian environments, even with hard state aggregation, using a simple action-commitment strategy.