collaborators

6 papers

cs.CL2026

Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis

Souranil Kahali, Rituparna Bose, Abner Hernandez +4

Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pret…

cs.CL2026

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

Abner Hernandez, Tomás Arias Vergara, Daiqi Liu +2

Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challengi…

cs.CL2026

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

Abner Hernandez, Tomás Arias-Vergara, Daiqi Liu +2

Phonological features provide a language-general and linguistically grounded representation of speech. We present PhonoQ-2.0, a multilingual frame-level phonological feature recogn…

cs.CV2026

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

Daiqi Liu, Lukas Mulzer, Md Hasan +11

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial re…

cs.SD2026

SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

Md Hasan, Nyvenn Castro, Daiqi Liu +6

Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for speech science and clinical a…

eess.AS2026

Perceptual implications of automatic anonymization in pathological speech

Soroosh Tayebi Arasteh, Saba Afza, Tri-Thien Nguyen +11

Automatic anonymization is increasingly used to enable ethical sharing of clinical speech, yet its perceptual and clinical consequences remain undercharacterized. We present a huma…