Search papers, labs, and topics across Lattice.
1
0
3
8
Reward-driven optimization in continuous-time RL can significantly enhance the fine-tuning of discrete diffusion models, even with non-differentiable reward signals.