3 papers
cs.SD2024
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
Li Zhang, Ning Jiang, Qing Wang +3
Trained on 680,000 hours of massive speech data, Whisper is a multitasking, multilingual speech foundation model demonstrating superior performance in automatic speech recognition,…
cs.SD2024
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
Kun Wei, Bei Li, Hang Lv +3
Automatic Speech Recognition (ASR) in conversational settings presents unique challenges, including extracting relevant contextual information from previous conversational turns. D…
eess.AS2024
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
Xinfa Zhu, Yuke Li, Yi Lei +3
This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastiv…