speech processing

Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?

arXiv:2607.13864

summary

The paper investigates how supervised fine-tuning performance for speech foundation models varies across different pretrained checkpoints, showing that gains often depend on the specific checkpoint and seed rather than universally better methods.

Abstract

Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained checkpoint are often interpreted as method-level improvements, i.e., a higher attainable performance ceiling. We show that such conclusions are not always reliable because SFT outcomes depend strongly on the specific pretrained instance. We conduct a systematic study on 3 SUPERB classification tasks, evaluating 8 SFT variants across 9 pretrained checkpoints from wav2vec~2.0, HuBERT, and WavLM, with multi-seed repetitions on representative base-scale models. We find that the identity of the statistically indistinguishable top-group SFT recipe is often checkpoint-dependent, with limited transferability across pretrained instances. These findings suggest that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling.

Accept by Interspeech 2026

Topics & keywords