checkpoint variability 1self-supervised speech models 1speech classification 1SUPERB evaluation 1supervised fine-tuning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.SD2026
Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
Wangjin Zhou, Yizhou Zhang, Yichi Wang +1
The paper investigates how supervised fine-tuning performance for speech foundation models varies across different pretrained checkpoints, showing that gains often depend on the sp…
cs.SD2026
Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR
Yichi Wang, Junzhe Chen, Wangjin Zhou +1
In long-form multi-party conversations, highly imbalanced speaker activity and frequent overlap make it difficult to identify "who spoke when and what". Sliding-window continuous s…
cs.SD2026
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
Yizhou Zhang, Yuan Gao, Wangjin Zhou +3
Self-supervised learning (SSL) on large-scale datasets like AudioSet has become the dominant paradigm for audio representation learning. While the continuous influx of new, unlabel…