Search papers, labs, and topics across Lattice.
Peking University
1
0
2
Enhancing on-policy reinforcement learning can paradoxically reduce the diversity of successful behaviors, leading to a trade-off that challenges future trainability.