Search papers, labs, and topics across Lattice.
This paper investigates the organization of content and speaker information in self-supervised speech features by analyzing the dimensions of WavLM-factorised subspaces. The authors find that leading dimensions in the content space are primarily associated with intensity and voicing, while pitch is encoded in later dimensions, and the highest-variance speaker dimension correlates strongly with pitch and gender. Their intervention experiments demonstrate that manipulating these dimensions allows for targeted control over speech characteristics, enhancing speech synthesis capabilities.
Manipulating specific dimensions in speech feature subspaces can enable precise control over characteristics like pitch and intensity, revolutionizing speech synthesis techniques.
Self-supervised speech features encode both content and speaker information. Recent work introduced an SVD-based factorisation that decomposes these features into a shared content matrix capturing temporal variation and speaker-specific transformations capturing static speaker characteristics. However, how information is organised within these components remains unclear. In this paper, we investigate how the dimensions of WavLM-factorised content and speaker subspaces correlate with speech characteristics such as pitch, intensity, and voicing. We find that leading dimensions in the content space primarily capture intensity, higher-order formants, and voicing, while pitch is encoded in a later dimension. In contrast, the highest-variance speaker dimension is strongly associated with pitch and gender, with later dimensions capturing high-frequency variation. Intervention experiments show that manipulating these dimensions enables targeted control of speech characteristics for speech synthesis. Furthermore, modifying the content and speaker representations jointly provides fine-grained control over characteristics such as pitch and intensity.