Search papers, labs, and topics across Lattice.
University of Toronto
1
0
2
UBP2 achieves a remarkable boost in sample efficiency for preference-based RL by actively balancing exploration and exploitation through uncertainty reasoning.