activity
20162023
most citedVery Deep Convolutional Neural Networks for Robust Speech Recognition

12 citations · 14 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CL2023

Can Contextual Biasing Remain Effective with Whisper and GPT-2?

Guangzhi Sun, Xianrui Zheng, Chao Zhang +1

End-to-end automatic speech recognition (ASR) and large language models, such as Whisper and GPT-2, have recently been scaled to use vast amounts of training data. Despite the larg…

cs.CL20231 cited

Graph Neural Networks for Contextual ASR with the Tree-Constrained Pointer Generator

Guangzhi Sun, Chao Zhang, Phil Woodland

The incorporation of biasing words obtained through contextual knowledge is of paramount importance in automatic speech recognition (ASR) applications. This paper proposes an innov…

eess.AS20231 cited

Self-Supervised Learning-Based Source Separation for Meeting Data

Yuang Li, Xianrui Zheng, Philip C. Woodland

Source separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the succes…

eess.AS20231 cited

Knowledge Distillation from Multiple Foundation Models for End-to-End Speech Recognition

Xiaoyu Yang, Qiujia Li, Chao Zhang +1

Although large foundation models pre-trained by self-supervised learning have achieved state-of-the-art performance in many tasks including automatic speech recognition (ASR), know…

eess.AS20231 cited

Adaptable End-to-End ASR Models using Replaceable Internal LMs and Residual Softmax

Keqi Deng, Philip C. Woodland

End-to-end (E2E) automatic speech recognition (ASR) implicitly learns the token sequence distribution of paired audio-transcript training data. However, it still suffers from domai…

eess.AS2022

Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription

Xianrui Zheng, Chao Zhang, Philip C. Woodland

Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diar…