4 papers
Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models
Yizhou Zhang, Wangjin Zhou, Xin Gu +6
Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent…
Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
Wangjin Zhou, Yizhou Zhang, Yichi Wang +1
Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained check…
Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR
Yichi Wang, Junzhe Chen, Wangjin Zhou +1
In long-form multi-party conversations, highly imbalanced speaker activity and frequent overlap make it difficult to identify "who spoke when and what". Sliding-window continuous s…
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
Yizhou Zhang, Yuan Gao, Wangjin Zhou +3
Self-supervised learning (SSL) on large-scale datasets like AudioSet has become the dominant paradigm for audio representation learning. While the continuous influx of new, unlabel…