93 citations · 224 across the 23 of their papers we have counts for
32 papers
Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using -VAE
Hui Lu, Disong Wang, Xixin Wu +3
We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot c…
Audio-visual multi-channel speech separation, dereverberation and recognition
Guinan Li, Jianwei Yu, Jiajun Deng +2
Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speak…
Neural Architecture Search for Speech Emotion Recognition
Xixin Wu, Shoukang Hu, Zhiyong Wu +2
Deep neural networks have brought significant advancements to speech emotion recognition (SER). However, the architecture design in SER is mainly based on expert knowledge and empi…
Exploiting Cross Domain Acoustic-to-articulatory Inverted Features For Disordered Speech Recognition
Shujie Hu, Shansong Liu, Xurong Xie +6
Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal spee…
Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition
Mengzhe Geng, Xurong Xie, Zi Ye +5
Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech in recent decades, accurate recognition of dysarthric and elderly speech remai…
Recent Progress in the CUHK Dysarthric Speech Recognition System
Shansong Liu, Mengzhe Geng, Shoukang Hu +5
Despite the rapid progress of automatic speech recognition (ASR) technologies in the past few decades, recognition of disordered speech remains a highly challenging task to date. D…