2 papers
cs.SD2021
Using multiple reference audios and style embedding constraints for speech synthesis
Cheng Gong, Longbiao Wang, Zhenhua Ling +2
The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the r…
cs.CV2020
Three-Dimensional Lip Motion Network for Text-Independent Speaker Recognition
Jianrong Wang, Tong Wu, Shanyu Wang +4
Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensi…