5 citations · 7 across the 6 of their papers we have counts for
11 papers
Improving the fusion of acoustic and text representations in RNN-T
Chao Zhang, Bo Li, Zhiyun Lu +2
The recurrent neural network transducer (RNN-T) has recently become the mainstream end-to-end approach for streaming automatic speech recognition (ASR). To estimate the output dist…
Adapting GPT, GPT-2 and BERT Language Models for Speech Recognition
Xianrui Zheng, Chao Zhang, Philip C. Woodland
Language models (LMs) pre-trained on massive amounts of text, in particular bidirectional encoder representations from Transformers (BERT), generative pre-training (GPT), and GPT-2…
Tree-constrained Pointer Generator for End-to-end Contextual Speech Recognition
Guangzhi Sun, Chao Zhang, Philip C. Woodland
Contextual knowledge is important for real-world automatic speech recognition (ASR) applications. In this paper, a novel tree-constrained pointer generator (TCPGen) component is pr…
Combining Frame-Synchronous and Label-Synchronous Systems for Speech Recognition
Qiujia Li, Chao Zhang, Philip C. Woodland
Commonly used automatic speech recognition (ASR) systems can be classified into frame-synchronous and label-synchronous categories, based on whether the speech is decoded on a per-…
A Distributed Optimisation Framework Combining Natural Gradient with Hessian-Free for Discriminative Sequence Training
Adnan Haider, Chao Zhang, Florian L. Kreyssig +1
This paper presents a novel natural gradient and Hessian-free (NGHF) optimisation framework for neural network training that can operate efficiently in a distributed manner. It rel…
Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis
Guanghui Xu, Wei Song, Zhengchen Zhang +3
Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which ma…