Search papers, labs, and topics across Lattice.
University of California, Los Angeles
1
0
Static token credit misaligns with evolving training dynamics, but Se-DPO's adaptive approach boosts performance by nearly 10 points on key benchmarks.