2 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.SD2024
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
Zehai Tu, Guangyan Zhang, Yiting Lu +3
Tokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Altho…
eess.AS2023★ 1 cited
Intelligibility prediction with a pretrained noise-robust automatic speech recognition model
Zehai Tu, Ning Ma, Jon Barker
This paper describes two intelligibility prediction systems derived from a pretrained noise-robust automatic speech recognition (ASR) model for the second Clarity Prediction Challe…
cs.SD2023★ 2 cited
Energy-Based Models For Speech Synthesis
Wanli Sun, Zehai Tu, Anton Ragni
Recently there has been a lot of interest in non-autoregressive (non-AR) models for speech synthesis, such as FastSpeech 2 and diffusion models. Unlike AR models, these models do n…