Search papers, labs, and topics across Lattice.
This study investigates preference optimization techniques for synthesizing non-verbal vocalizations (NVs) in text-to-speech (TTS) systems, addressing a gap in understanding their effectiveness. By developing a novel NV-aware character error rate (NV-CER) metric, the authors evaluate the impact of various design choices on NV realization and lexical fidelity across multiple NV types. The results demonstrate that standard decision-focused optimization (DPO) can effectively enhance NV generation without altering the core optimization framework, validated through objective and human evaluations.
NV synthesis can be optimized without changing the underlying algorithms, revealing critical insights into TTS expressiveness.
Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for NV-capable TTS, focusing on preference signals, preference-pair construction, and DPO-based optimization objectives. We formulate an NV-aware character error rate (NV-CER) by treating NV tags as distinct output symbols and computing a weighted pinyin-based CER over both verbal and non-verbal content, enabling controllable optimization of NV realization without modifying the underlying optimization algorithm. Experiments on Emilia-NV and the augmented NV-Bench covering 18 NV types reveal how different design choices affect NV realization and lexical fidelity, and establish an effective setup using standard DPO. Objective, LLM-based, and human evaluations provide converging evidence for our findings, offering practical insights into NV-aware post-training for expressive TTS.