Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
0
A two-stage OPD-then-RL approach outperforms traditional methods by leveraging the strengths of both on-policy distillation and reinforcement learning without the interference seen in joint optimization.