activity
20192025
most citedGenerating diverse and natural text-to-speech samples using a quantized fine-grained VAE and auto-regressive prosody prior

15 citations · 35 across the 9 of their papers we have counts for

collaborators

8 papers

eess.AS2024

SOT Triggered Neural Clustering for Speaker Attributed ASR

Xianrui Zheng, Guangzhi Sun, Chao Zhang +1

This paper introduces a novel approach to speaker-attributed ASR transcription using a neural clustering method. With a parallel processing mechanism, diarisation and ASR can be ap…

cs.SD2022

Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech

Yang Li, Cheng Yu, Guangzhi Sun +6

Modelling prosody variation is critical for synthesizing natural and expressive speech in end-to-end text-to-speech (TTS) systems. In this paper, a cross-utterance conditional VAE…

cs.CL20212 cited

Tree-constrained Pointer Generator for End-to-end Contextual Speech Recognition

Guangzhi Sun, Chao Zhang, Philip C. Woodland

Contextual knowledge is important for real-world automatic speech recognition (ASR) applications. In this paper, a novel tree-constrained pointer generator (TCPGen) component is pr…

cs.SD20212 cited

Content-Aware Speaker Embeddings for Speaker Diarisation

G. Sun, D. Liu, C. Zhang +1

Recent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddi…

cs.CL20205 cited

Cross-Utterance Language Models with Acoustic Error Sampling

G. Sun, C. Zhang, P. C. Woodland

The effective exploitation of richer contextual information in language models (LMs) is a long-standing research problem for automatic speech recognition (ASR). A cross-utterance L…

eess.AS202015 cited

Generating diverse and natural text-to-speech samples using a quantized fine-grained VAE and auto-regressive prosody prior

Guangzhi Sun, Yu Zhang, Ron J. Weiss +5

Recent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-gr…