Search papers, labs, and topics across Lattice.
Beijing Jiaotong University (Weihai)
1
0
2
2
This paper investigates the development path of large language model preference alignment techniques, focusing on the training mechanism of Reinforcement Learning from Human Feedback and its three-stage process, including supervised fine-tuning, reward model training, and PPO-based policy optimization.