Search papers, labs, and topics across Lattice.
2
0
5
5
Diffusion models can navigate more reliably without retraining, thanks to a clever guidance method that keeps them on the training manifold.
Forget hand-crafted reward functions: $\text{RLR}^3$ leverages rubrics and LLMs to provide fine-grained, multi-criteria supervision, outperforming standard RLVR in vision-language tasks.