3 papers
eess.AS2025
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
Dong Liu, Juan Liu, Wei Ju +2
Whispered speech lacks vocal-fold excitation, making intelligible conversion challenging. We propose WhisperVC, a three-stage framework for low-resource whisper-to-normal (W2N) con…
cs.SD2024
Voice EHR: Introducing Multimodal Audio Data for Health
James Anibal, Hannah Huth, Ming Li +27
Artificial intelligence (AI) models trained on audio data may have the potential to rapidly perform clinical tasks, enhancing medical decision-making and potentially improving outc…
cs.SD2023
Pretraining Conformer with ASR or ASV for Anti-Spoofing Countermeasure
Yikang Wang, Hiromitsu Nishizaki, Ming Li
Finding synthetic artifacts of spoofing data will help the anti-spoofing countermeasures (CMs) system discriminate between spoofed and real speech. The Conformer combines the best…