Search papers, labs, and topics across Lattice.
Hanyang University
1
0
3
DPO-tuned LLMs systematically achieve less than a third of requested emotional intensity because preference optimization cannot reward behavioral extremes that base models never sample.