activity
20242026
collaborators

5 papers

cs.SD2026

Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

Mattias Cross, Minghui Zhao, Anton Ragni

Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this leng…

cs.LG2025

Discrete-Time Diffusion-Like Models for Speech Synthesis

Xiaozhou Tan, Minghui Zhao, Anton Ragni

Diffusion models have attracted a lot of attention in recent years. These models view speech generation as a continuous-time process. For efficient training, this process is typica…

cs.SD2025

Flowing Straighter with Conditional Flow Matching for Accurate Speech Enhancement

Mattias Cross, Anton Ragni

Current flow-based generative speech enhancement methods learn curved probability paths which model a mapping between clean and noisy speech. Despite impressive performance, the im…

cs.LG2024

What happens to diffusion model likelihood when your model is conditional?

Mattias Cross, Anton Ragni

Diffusion Models (DMs) iteratively denoise random samples to produce high-quality data. The iterative sampling process is derived from Stochastic Differential Equations (SDEs), all…

cs.SD2024

Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis

Wing-Zin Leung, Mattias Cross, Anton Ragni +1

Automatic speech recognition (ASR) research has achieved impressive performance in recent years and has significant potential for enabling access for people with dysarthria (PwD) i…