Search papers, labs, and topics across Lattice.
University of Chinese Academy of Sciences
1
0
2
Sparse rewards can be effectively optimized without losing the advantages of fine-grained feedback through a novel two-stage training approach.