Search papers, labs, and topics across Lattice.
This study revisits the Vocos architecture for time-frequency neural vocoding, identifying its strengths in magnitude modeling but weaknesses in phase prediction. By quantifying the performance gap between time-domain and time-frequency vocoders using bandlimited mel spectrograms, the authors reveal that 1D convolutional layers impede accurate phase reconstruction. The findings suggest a need for future architectures to incorporate inductive biases that enhance modeling of the time-frequency structure of speech signals while maintaining flexibility for various input types.
Vocos excels at magnitude modeling but falters on phase, revealing critical architectural limitations that could redefine vocoding approaches.
Recently, time-frequency neural vocoders have been approaching the state-of-the-art quality of time-domain neural vocoders. Vocos is a notable example due to its efficiency, but its audio quality lags behind the time-domain vocoders and the reasons remain debated. Thus, in this study, we revisit Vocos from a phase reconstruction perspective. First, we quantify the gap between time-domain and time-frequency domain vocoders using bandlimited mel spectrograms as inputs. Later, via an ablation study, we verify the Vocos architecture is effective for magnitude modeling, but less so for phase. We then adapt the Vocos backbone to predict phase differences, a precursor for phase reconstruction, and identify 1D convolutional layers are hindering their accurate prediction. Our findings indicate that future research needs to focus on inductive biases that allow the architecture to better model the time-frequency structure of speech signals, without sacrificing the support for arbitrary input representations.