5 papers
Beyond Two-stage Diffusion TTS: Joint Structure and Content Refinement via Jump Diffusion
Jiabao Ai, Minghui Zhao, Anton Ragni
Diffusion and flow matching TTS faces a tension between discrete temporal structure and continuous spectral modeling. Two-stage models diffuse on fixed alignments, often collapsing…
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
Robert Flynn, Anton Ragni
Automatic speech recognition (ASR) models are normally trained to operate over single utterances, with a short duration of less than 30 seconds. This choice has been made in part d…
Score-Based Training for Energy-Based TTS Models
Wanli Sun, Anton Ragni
Noise contrastive estimation (NCE) is a popular method for training energy-based models (EBM) with intractable normalisation terms. The key idea of NCE is to learn by comparing unn…
Self-Train Before You Transcribe
Robert Flynn, Anton Ragni
When there is a mismatch between the training and test domains, current speech recognition systems show significant performance degradation. Self-training methods, such as noisy st…
How Much Context Does My Attention-Based ASR System Need?
Robert Flynn, Anton Ragni
For the task of speech recognition, the use of more than 30 seconds of acoustic context during training is uncommon and under-investigated in literature. In this work, we conduct a…