activity
20202022
most citedTowards Multi-Scale Style Control for Expressive Speech Synthesis

4 citations · 11 across the 7 of their papers we have counts for

collaborators

13 papers

cs.LG2022

Ordinal Regression via Binary Preference vs Simple Regression: Statistical and Experimental Perspectives

Bin Su, Shaoguang Mao, Frank Soong +1

Ordinal regression with anchored reference samples (ORARS) has been proposed for predicting the subjective Mean Opinion Score (MOS) of input stimuli automatically. The ORARS addres…

cs.MM2021

Transformer-S2A: Robust and Efficient Speech-to-Animation

Liyang Chen, Zhiyong Wu, Jun Ling +3

We propose a novel robust and efficient Speech-to-Animation (S2A) approach for synchronized facial animation generation in human-computer interaction. Compared with conventional ap…

cs.CL2021

An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings

Wenxuan Ye, Shaoguang Mao, Frank Soong +4

Many mispronunciation detection and diagnosis (MD&D) research approaches try to exploit both the acoustic and linguistic features as input. Yet the improvement of the performance i…

cs.CL2021

Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding

Yingmei Guo, Linjun Shou, Jian Pei +4

Lack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages. Although various data augmentation approaches have be…

cs.LG2021★ 3 cited

Voting for the right answer: Adversarial defense for speaker verification

Haibin Wu, Yang Zhang, Zhiyong Wu +2

Automatic speaker verification (ASV) is a well developed technology for biometric identification, and has been ubiquitous implemented in security-critic applications, such as banki…

cs.SD2021

VAENAR-TTS: Variational Auto-Encoder based Non-AutoRegressive Text-to-Speech Synthesis

Hui Lu, Zhiyong Wu, Xixin Wu +4

This paper describes a variational auto-encoder based non-autoregressive text-to-speech (VAENAR-TTS) model. The autoregressive TTS (AR-TTS) models based on the sequence-to-sequence…