Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
Visual signals, not language priors, drive attribute hallucination in VLMs, leading to a novel framework that effectively mitigates this issue.
Deterministic policies can significantly enhance the stability and efficiency of reinforcement learning in complex mean field control problems, outperforming traditional stochastic approaches.
WPO, a promising RL algorithm for continuous control, is now proven to converge linearly, finally putting it on solid theoretical footing.
Asynchronous RL for LLMs doesn't have to sacrifice convergence for speed: DORA achieves 2-4x faster training by cleverly managing multiple policy versions during rollout.