6 papers
Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis
Souranil Kahali, Rituparna Bose, Abner Hernandez +4
Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pret…
Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification
Abner Hernandez, Tomás Arias Vergara, Daiqi Liu +2
Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challengi…
Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
Abner Hernandez, Tomás Arias-Vergara, Daiqi Liu +2
Phonological features provide a language-general and linguistically grounded representation of speech. We present PhonoQ-2.0, a multilingual frame-level phonological feature recogn…
Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI
Daiqi Liu, Lukas Mulzer, Md Hasan +11
Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial re…
SIREM: Speech-Informed MRI Reconstruction with Learned Sampling
Md Hasan, Nyvenn Castro, Daiqi Liu +6
Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for speech science and clinical a…
Perceptual implications of automatic anonymization in pathological speech
Soroosh Tayebi Arasteh, Saba Afza, Tri-Thien Nguyen +11
Automatic anonymization is increasingly used to enable ethical sharing of clinical speech, yet its perceptual and clinical consequences remain undercharacterized. We present a huma…