5 papers
Decoding Order Matters in Autoregressive Speech Synthesis
Minghui Zhao, Anton Ragni
Autoregressive speech synthesis often adopts a left-to-right order, yet generation order is a modelling choice. We investigate decoding order through masked diffusion framework, wh…
A comparative study of generative models for child voice conversion
Protima Nomo Sudro, Anton Ragni, Thomas Hain
Generative models are a popular choice for adult-to-adult voice conversion (VC) because of their efficient way of modelling unlabelled data. To this point their usefulness in produ…
Discrete-Time Diffusion-Like Models for Speech Synthesis
Xiaozhou Tan, Minghui Zhao, Anton Ragni
Diffusion models have attracted a lot of attention in recent years. These models view speech generation as a continuous-time process. For efficient training, this process is typica…
Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement
Mattias Cross, Anton Ragni
Current flow-based generative speech enhancement methods learn curved probability paths which model a mapping between clean and noisy speech. Despite impressive performance, the im…
Emphasis Sensitivity in Speech Representations
Shaun Cassini, Thomas Hain, Anton Ragni
This work investigates whether modern speech models are sensitive to prosodic emphasis - whether they encode emphasized and neutral words in systematically different ways. Prior wo…