3 papers
eess.AS2025
What Does the Speaker Embedding Encode?
Shuai Wang, Yanmin Qian, Kai Yu
Developing a good speaker embedding has received tremendous interest in the speech community, with representations such as i-vector and d-vector demonstrating remarkable performanc…
eess.AS2025
Text adaptation for speaker verification with speaker-text factorized embeddings
Yexin Yang, Shuai Wang, Xun Gong +2
Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system p…
cs.SD2025
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang +1
We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel mu…