Search papers, labs, and topics across Lattice.
To resolve the severe over-refusal and utility degradation plaguing safety alignment in open-weight LLMs, the authors develop Suan, a direct preference optimization framework formulated explicitly at the gradient level. Bypassing traditional variational loss derivations enables finer, more interpretable control over policy updates during safety tuning. Across diverse safety and capability benchmarks, the method achieves robust guardrail compliance while fully maintaining model helpfulness and general performance.
Formulating preference optimization directly at the gradient level eliminates the notorious over-refusal penalty in open-weight LLMs without compromising downstream utility.
Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynamics. Extensive evaluations across a diverse suite of competitive baselines and benchmarks demonstrate that Suan achieves superior safety alignment while fully preserving response utility.