paper

Controllable Dysarthric Speech Synthesis with Patient-Specific Conditioning for Speaker-Diverse ASR Augmentation

arXiv:2602.08696

Abstract

Dysarthric speech recognition is limited by high speaker variability and scarce labeled data. Existing synthesis methods often couple speaker identity with dysarthric articulation, reducing control over generated speech. We propose a controllable dysarthric speech synthesis framework for ASR augmentation with separate prompt-derived timbre prefixes and learnable patient-specific pathology prefixes. Built on a pre-trained neural codec language model adapted using LoRA, the framework combines both prefixes through additive conditioning. A dual-classifier objective with gradient reversal and voice-conversion-based counterfactual augmentation promotes factor separation, allowing learned pathology conditions to be paired with different target speakers, including healthy speakers. Experiments on TORGO show that generated speech can partially replace real dysarthric training data and provides effective speaker-diverse augmentation when combined with real data. Objective ASR, perceptual, factor-separation, and phoneme-level analyses indicate preservation of target-speaker timbre and pathology-dependent patterns consistent with real dysarthric speech.

Controllable Dysarthric Speech Synthesis with Patient-Specific Conditioning for Speaker-Diverse ASR Augmentation · wovepaper