Search papers, labs, and topics across Lattice.
The University of Hong Kong, Hong Kong SAR
2
0
3
Regularizing the SAIL objective with reverse KL divergence not only resolves convergence issues but also enhances performance in LLM alignment tasks.
Extracting temporal geometry from generative models can boost reinforcement learning performance by over 2x without changing the optimal policy.