5 papers
What Does the Speaker Embedding Encode?
Shuai Wang, Yanmin Qian, Kai Yu
Developing a good speaker embedding has received tremendous interest in the speech community, with representations such as i-vector and d-vector demonstrating remarkable performanc…
Text adaptation for speaker verification with speaker-text factorized embeddings
Yexin Yang, Shuai Wang, Xun Gong +2
Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system p…
Data Augmentation for End-to-end Code-switching Speech Recognition
Chenpeng Du, Hao Li, Yizhou Lu +2
Training a code-switching end-to-end automatic speech recognition (ASR) model normally requires a large amount of data, while code-switching data is often limited. In this paper, t…
DDTSE: Discriminative Diffusion Model for Target Speech Extraction
Leying Zhang, Yao Qian, Linfeng Yu +5
Diffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction u…
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
Hang Shao, Bei Liu, Wei Wang +2
As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models,…