Search papers, labs, and topics across Lattice.
The paper introduces CLASVS, a method for melody-preserving lyric editing in singing voice synthesis that utilizes continuous-latent autoregression to enhance the editing process. By employing State-Control-Transition (SCT) routing and Progressive State-Control Grounding (PSCG), CLASVS effectively maintains the integrity of melody and singer identity while allowing for semantic feedback during lyric modifications. The results demonstrate a 46.2% reduction in macro-PER compared to discrete-AR models, showcasing significant improvements in lyric editing without compromising audio quality or performance characteristics.
CLASVS achieves a remarkable 46.2% reduction in editing errors while preserving melody and singer identity, revolutionizing lyric editing in singing voice synthesis.
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict absent from ordinary reconstruction: training pairs reference cues with original lyrics, whereas inference asks revised lyrics to override source-lyric-correlated cues; one source-following patch can propagate through AR history. We introduce CLASVS. Its State-Control-Transition (SCT) routing keeps target-lyric and reference-melody controls persistent, returns semantic feedback on phonetic progress to the causal planner, and confines the previous latent patch to the local Transition. Progressive State-Control Grounding (PSCG) learns this contract through paired-edit-free, content-consistent Mandarin reconstruction. On two Mandarin benchmarks, CLASVS improves all four operations over discrete-AR Vevo2 and reduces macro-PER by 46.2%, while maintaining melody, singer similarity, and perceptual quality. Together, these results establish a strong continuous-AR operating point for score-annotation-free lyric edits and a basis for broader stepwise control. Audio demonstrations are available on our project page: https://piedpiperg.github.io/Liyric-SVS/.