activity
20192023
most citedRecent Progress in the CUHK Dysarthric Speech Recognition System

93 citations · 300 across the 32 of their papers we have counts for

collaborators
Showing 2020 · eess.ASShow all

6 papers · 2 filters

eess.AS2020

Bayesian Learning of LF-MMI Trained Time Delay Neural Networks for Speech Recognition

Shoukang Hu, Xurong Xie, Shansong Liu +5

Discriminative training techniques define state-of-the-art performance for automatic speech recognition systems. However, they are inherently prone to overfitting, leading to poor…

eess.AS2020★ 6 cited

Improved End-to-End Dysarthric Speech Recognition via Meta-learning Based Model Re-initialization

Disong Wang, Jianwei Yu, Xixin Wu +3

Dysarthric speech recognition is a challenging task as dysarthric data is limited and its acoustics deviate significantly from normal speech. Model-based speaker adaptation is a pr…

eess.AS2020

Audio-visual Multi-channel Integration and Recognition of Overlapped Speech

Jianwei Yu, Shi-Xiong Zhang, Bo Wu +6

Automatic speech recognition (ASR) technologies have been significantly advanced in the past few decades. However, recognition of overlapped speech remains a highly challenging tas…

eess.AS2020

Audio-visual Multi-channel Recognition of Overlapped Speech

Jianwei Yu, Bo Wu, Rongzhi Gu +7

Automatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-…

eess.AS2020★ 6 cited

Bayesian x-vector: Bayesian Neural Network based x-vector System for Speaker Verification

Xu Li, Jinghua Zhong, Jianwei Yu +4

Speaker verification systems usually suffer from the mismatch problem between training and evaluation data, such as speaker population mismatch, the channel and environment variati…

eess.AS2020★ 10 cited

Audio-visual Recognition of Overlapped speech for the LRS2 dataset

Jianwei Yu, Shi-Xiong Zhang, Jian Wu +7

Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of…