Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
5
QDOS transforms offline datasets into a treasure trove of high-value skills, significantly enhancing reinforcement learning performance in challenging tasks.
Dynamically learning support bounds for value functions can enhance stability and performance in reinforcement learning, outperforming traditional fixed-interval approaches.
Taming policy diversity with KL constraints unlocks surprisingly stable and sample-efficient ensemble reinforcement learning in high-dimensional manipulation tasks.