5 papers
Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning
Roseline Polle, Owen Parsons, George Fairs +5
Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled…
Towards Dys-XAI: Influence-Based Explanations for Dysarthria Severity Assessment
Xiaoliang Wu, Qiyang Sun, Yupei Li +3
Dysarthria severity assessment is essential for therapy planning and longitudinal monitoring, yet manual perceptual rating is time-consuming and variable across clinicians. Althoug…
LISE : Listenable Interpretable Speaker Embeddings
Xiaoliang Wu, Chongxin Gan, Ke Liu +2
Deep neural network-based automatic speaker verification (ASV) systems achieve impressive performance but their embedding representations remain opaque, lacking a structured and pe…
Explanations for Automatic Speech Recognition
Xiaoliang Wu, Peter Bell, Ajitha Rajan
We address quality assessment for neural network based ASR by providing explanations that help increase our understanding of the system and ultimately help build trust in the syste…
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
Yupei Li, Qiyang Sun, Xiaoliang Wu +3
Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explai…