Search papers, labs, and topics across Lattice.
This paper introduces MeloCodec, a novel framework that integrates melodic priors into neural audio codecs to enhance the fidelity of singing voice representation. By employing a Tokenize-then-Fuse paradigm and a two-stage training strategy, the authors effectively mitigate optimization instability and prevent codebook collapse, which are common challenges in incorporating explicit acoustic information. Experimental results demonstrate that MeloCodec significantly improves pitch consistency and allows for controllable pitch manipulation with minimal degradation of timbre compared to existing methods.
MeloCodec achieves superior singing voice representation by leveraging melodic priors, enabling precise pitch control without compromising timbre quality.
Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.