Search papers, labs, and topics across Lattice.
Delft University of Technology
3
0
6
A strategic increase in policy sets can dramatically reduce regret in uncertain environments, challenging the conventional wisdom of single-policy optimization.
By learning to mask attention weights, SMAP enables reinforcement learning agents to generalize far better to unseen environments in Procgen.
Random Network Distillation, a computationally cheap uncertainty method, is theoretically equivalent to both deep ensembles and Bayesian inference under certain conditions, finally giving it a solid theoretical footing.