Search papers, labs, and topics across Lattice.
Southeast University, This work was supported by the State Key Laboratory of Autonomous Intelligent Unmanned Systems (ZZKF2025-2-5) and the State Key Laboratory of Robotics and Systems (HIT) (SKLRS-2025-KF-11). (Corresponding author: Longhui Qin)Hongliang Zhao, Wenhui Yang, Yang Chen, Zhuorui Wang, and Baiheng Liu are with the School of Mechanical Engineering, Southeast University, Nanjing 211189, China. E-mail: zhaohongliang@msn.com; wenhui-yang@outlook.com.Longhui Qin is with the School of Mechanical Engineering, Southeast University, Nanjing, 211189, China, the State Key Laboratory of Autonomous Intelligent Unmanned Systems, Beijing Institute of Technology, Beijing, 10081, China and the State Key Laboratory of Robotics and Systems, Harbin Institute of Technology, Harbin, 150001, China. (e-mail: lhqin@seu.edu.cn)
1
7
3
1
By recognizing that not all tokens are created equal, D2PO offers a simple temporal weighting fix that boosts DPO alignment scores by up to 9.7 points.