Search papers, labs, and topics across Lattice.
This study critically evaluates the effectiveness of supervised fine-tuning (SFT) for adapting self-supervised speech models to classification tasks, revealing that performance gains are often contingent on the specific pretrained model rather than inherent improvements in SFT methods. By systematically analyzing eight SFT variants across nine pretrained checkpoints from leading models like wav2vec~2.0, HuBERT, and WavLM on three SUPERB tasks, the authors demonstrate that the best-performing SFT approach varies significantly depending on the pretrained instance. This challenges the prevailing assumption that SFT universally enhances performance ceilings, suggesting that many observed gains may stem from instance-specific characteristics rather than methodological advancements.
Performance improvements in speech model fine-tuning are often an illusion, heavily reliant on the specific pretrained instance rather than true methodological advancements.
Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained checkpoint are often interpreted as method-level improvements, i.e., a higher attainable performance ceiling. We show that such conclusions are not always reliable because SFT outcomes depend strongly on the specific pretrained instance. We conduct a systematic study on 3 SUPERB classification tasks, evaluating 8 SFT variants across 9 pretrained checkpoints from wav2vec~2.0, HuBERT, and WavLM, with multi-seed repetitions on representative base-scale models. We find that the identity of the statistically indistinguishable top-group SFT recipe is often checkpoint-dependent, with limited transferability across pretrained instances. These findings suggest that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling.