activity
20202022
most citedTowards Multi-Scale Style Control for Expressive Speech Synthesis

4 citations · 7 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2022

Lexical Knowledge Internalization for Neural Dialog Generation

Zhiyong Wu, Wei Bi, Xiang Li +2

We propose knowledge internalization (KI), which aims to complement the lexical knowledge into neural dialog models. Instead of further conditioning the knowledge-grounded dialog (…

cs.CL2022

An End-to-end Chinese Text Normalization Model based on Rule-guided Flat-Lattice Transformer

Wenlin Dai, Changhe Song, Xiang Li +4

Text normalization, defined as a procedure transforming non standard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. R…

cs.CL20213 cited

Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation

Zhiyong Wu, Lingpeng Kong, Wei Bi +2

A neural multimodal machine translation (MMT) system is one that aims to perform better translation by extending conventional text-only translation models with multimodal informati…

cs.SD20214 cited

Towards Multi-Scale Style Control for Expressive Speech Synthesis

Xiang Li, Changhe Song, Jingbei Li +3

This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract…

eess.AS2020

Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition

Xiong Cai, Dongyang Dai, Zhiyong Wu +3

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In…