1.6k citations
- Chinese Academy of SciencesCN17 papers
- Tsinghua UniversityCN17 papers
- Ministry of Industry and Information TechnologyCN16 papers
- Xi'an Jiaotong UniversityCN16 papers
- Xidian UniversityCN16 papers
- Australian National UniversityAU11 papers
- Moscow Institute of Physics and TechnologyRU11 papers
- University of Chinese Academy of SciencesCN10 papers
- Beijing Normal UniversityCN9 papers
- Nankai UniversityCN9 papers
- Centre National de la Recherche ScientifiqueFR8 papers
- Fudan UniversityCN8 papers
7 papers · 2 filters
Context-aware RNNLM Rescoring for Conversational Speech Recognition
Kun Wei, Pengcheng Guo, Hang Lv +2
Conversational speech recognition is regarded as a challenging task due to its free-style speaking and long-term contextual dependencies. Prior work has explored the modeling of lo…
Controllable Emotion Transfer For End-to-End Speech Synthesis
Tao Li, Shan Yang, Liumeng Xue +1
Emotion embedding space learned from references is a straightforward approach for emotion transfer in encoder-decoder structured emotional text to speech (TTS) systems. However, th…
Adversarial Training for Multi-domain Speaker Recognition
Qing Wang, Wei Rao, Pengcheng Guo +1
In real-life applications, the performance of speaker recognition systems always degrades when there is a mismatch between training and evaluation data. Many domain adaptation meth…
Accent and Speaker Disentanglement in Many-to-many Voice Conversion
Zhichao Wang, Wenshuo Ge, Xiong Wang +6
This paper proposes an interesting voice and accent joint conversion approach, which can convert an arbitrary source speaker's voice to a target speaker with non-native accent. Thi…
Fine-grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis
Yi Lei, Shan Yang, Lei Xie
This paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotion…
Cascade RNN-Transducer: Syllable Based Streaming On-device Mandarin Speech Recognition with a Syllable-to-Character Converter
Xiong Wang, Zhuoyuan Yao, Xian Shi +1
End-to-end models are favored in automatic speech recognition (ASR) because of its simplified system structure and superior performance. Among these models, recurrent neural networ…