93 citations · 178 across the 24 of their papers we have counts for
7 papers
Speech Enhancement with Fullband-Subband Cross-Attention Network
Jun Chen, Wei Rao, Zilin Wang +5
FullSubNet has shown its promising performance on speech enhancement by utilizing both fullband and subband information. However, the relationship between fullband and subband in F…
Towards High-Quality Neural TTS for Low-Resource Languages by Learning Compact Speech Representations
Haohan Guo, Fenglong Xie, Xixin Wu +2
This paper aims to enhance low-resource TTS by reducing training data requirements using compact speech representations. A Multi-Stage Multi-Codebook (MSMC) VQ-GAN is trained to le…
Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using -VAE
Hui Lu, Disong Wang, Xixin Wu +3
We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot c…
Robust Unsupervised Cross-Lingual Word Embedding using Domain Flow Interpolation
Liping Tang, Zhen Li, Zhiquan Luo +1
This paper investigates an unsupervised approach towards deriving a universal, cross-lingual word embedding space, where words with similar semantics from different languages are c…
Push-Pull: Characterizing the Adversarial Robustness for Audio-Visual Active Speaker Detection
Xuanjun Chen, Haibin Wu, Helen Meng +2
Audio-visual active speaker detection (AVASD) is well-developed, and now is an indispensable front-end for several multi-modal applications. However, to the best of our knowledge,…
Neural Architecture Search For LF-MMI Trained Time Delay Neural Networks
Shoukang Hu, Xurong Xie, Mingyu Cui +6
State-of-the-art automatic speech recognition (ASR) system development is data and computation intensive. The optimal design of deep neural networks (DNNs) for these systems often…