Search papers, labs, and topics across Lattice.
2
0
3
UBP2 achieves a remarkable boost in sample efficiency for preference-based RL by actively balancing exploration and exploitation through uncertainty reasoning.
Robots can now learn to balance waiting and rerouting in real-time, achieving near-optimal navigation even in cluttered environments.