Search papers, labs, and topics across Lattice.
1
0
2
Flow-matching models can now be fine-tuned with GRPO using only a third of the training steps, thanks to a novel off-policy approach.