Search papers, labs, and topics across Lattice.
This study challenges the conventional wisdom that increasing data volume improves contrastive audio embeddings, revealing that structural aspects of the corpus are more critical. By introducing a lexical-speech round to a frozen-base multimodal embedding model, the authors achieved a remarkable 76-point increase in zero-shot keyword spotting, albeit with a 14-point decrease in speech-emotion recognition. Fine-tuning on a carefully curated prosody-controlled corpus demonstrated that the loss in emotion recognition is not due to capacity limits but rather the nature of the data structure, emphasizing that the separability of attributes in contrastive learning is influenced more by corpus design than sheer volume.
Structural design of the corpus, not its size, dictates what attributes a contrastive audio embedding can effectively encode.
Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding model raises zero-shot keyword spotting by 76 points while reducing speech-emotion recognition by 14. The loss is not a capacity limit: fine-tuning on 7,442 clips from a prosody-controlled corpus recovers emotion past its pre-speech level at a five-point keyword cost. Nor is it data volume: 29,428 mined clips whose captions explicitly name emotions, at matched exposure, move emotion by -0.0007. The difference is structural: a contrastive objective encodes an attribute only when the in-batch negatives cannot be separated without it; the controlled corpus holds sentence content fixed, so prosody is the only separating signal, whereas mined captions name emotion yet remain separable by scene content. Intervention on the same audio confirms causality: raising caption similarity does not recover emotion, but collapsing caption diversity so that emotion becomes the only separating axis recovers it by 8.9 points across three seeds, with a smaller, same-signed gain on a non-acted corpus, while keyword accuracy trades back. Corpus structure, not size or caption vocabulary, controls what a contrastive audio embedding encodes.