Search papers, labs, and topics across Lattice.
Hanyang University
2
0
5
DPO-tuned LLMs systematically achieve less than a third of requested emotional intensity because preference optimization cannot reward behavioral extremes that base models never sample.
Vetoing unstable tokens can boost dMLLM reasoning accuracy by up to 9% without any extra training costs.