93 citations · 213 across the 14 of their papers we have counts for
18 papers
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Jianwei Yu +3
Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…
Music Source Separation with Band-split RNN
Yi Luo, Jianwei Yu
The performance of music source separation (MSS) models has been greatly improved in recent years thanks to the development of novel neural network architectures and training pipel…
Audio-visual multi-channel speech separation, dereverberation and recognition
Guinan Li, Jianwei Yu, Jiajun Deng +2
Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speak…
Improving Target Sound Extraction with Timestamp Information
Helin Wang, Dongchao Yang, Chao Weng +2
Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the p…
Recent Progress in the CUHK Dysarthric Speech Recognition System
Shansong Liu, Mengzhe Geng, Shoukang Hu +5
Despite the rapid progress of automatic speech recognition (ASR) technologies in the past few decades, recognition of disordered speech remains a highly challenging task to date. D…
Investigation of Data Augmentation Techniques for Disordered Speech Recognition
Mengzhe Geng, Xurong Xie, Shansong Liu +4
Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disab…