7 papers
Towards Dys-XAI: Influence-Based Explanations for Dysarthria Severity Assessment
Xiaoliang Wu, Qiyang Sun, Yupei Li +3
Dysarthria severity assessment is essential for therapy planning and longitudinal monitoring, yet manual perceptual rating is time-consuming and variable across clinicians. Althoug…
Phonetic Error Analysis of Raw Waveform Acoustic Models
Erfan Loweimi, Zhengjun Yue, Andrea Carmantini +3
We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER). PER is decomposed across three broad phonetic class…
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection
Erfan Loweimi, Mengjie Qian, Kate Knill +7
When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may b…
Predicting Psychological Well-Being from Spontaneous Speech using LLMs
Erfan Loweimi, Sofia de la Fuente Garcia, Saturnino Luz
We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. Using a few minutes of voice…
Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
Erfan Loweimi, Sofia de la Fuente Garcia, Samira Loveymi +2
LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability across three dimensions: intra…
Probing Whisper for Dysarthric Speech in Detection and Assessment
Zhengjun Yue, Devendra Kayande, Zoran Cvetkovic +1
Large-scale end-to-end models such as Whisper have shown strong performance on diverse speech tasks, but their internal behavior on pathological speech remains poorly understood. U…