WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
arXiv:2406.18731 · doi:10.1109/JBHI.2024.3454550
Abstract
Speech is known to carry health-related attributes, which has emerged as a novel venue for remote and long-term health monitoring. However, existing models are usually tailored for a specific type of disease, and have been shown to lack generalizability across datasets. Furthermore, concerns have been raised recently towards the leakage of speaker identity from health embeddings. To mitigate these limitations, we propose WavRx, a speech health diagnostics model that captures the respiration and articulation related dynamics from a universal speech representation. Our in-domain and cross-domain experiments on six pathological speech datasets demonstrate WavRx as a new state-of-the-art health diagnostic model. Furthermore, we show that the amount of speaker identity entailed in the WavRx health embeddings is significantly reduced without extra guidance during training. An in-depth analysis of the model was performed, thus providing physiological interpretation of its improved generalizability and privacy-preserving ability.
Under review; Model script available at https://github.com/zhu00121/WavRx
References in corpus (11)
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- X-vectors: New Quantitative Biomarkers for Early Parkinson's Disease Detection from Speech
- Wav2vec-based Detection and Severity Level Classification of Dysarthria from Speech
- A Noise-Robust Self-supervised Pre-training Model Based Speech Representation Learning for Automatic Speech Recognition
- Exploring Self-Supervised Representation Ensembles for COVID-19 Cough Classification
- A Step Towards Preserving Speakers' Identity While Detecting Depression Via Speaker Disentanglement
- Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals
- Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
- On the Impact of Voice Anonymization on Speech Diagnostic Applications: a Case Study on COVID-19 Detection