activity
20242026
collaborators

5 papers

eess.AS2026

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

Jinghao Chen, Mostafa Shahin, Beena Ahmed

Automatic mispronunciation detection and diagnosis (MDD) plays a crucial role in L2 Mandarin pronunciation learning. While end-to-end (E2E) based MDD methods have substantially imp…

cs.HC2025

System X: A Mobile Voice-Based AI System for EMR Generation and Clinical Decision Support in Low-Resource Maternal Healthcare

Maryam Mustafa, Umme Ammara, Amna Shahnawaz +5

We present the design, implementation, and in-situ deployment of a smartphone-based voice-enabled AI system for generating electronic medical records (EMRs) and clinical risk alert…

cs.SD2025

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM

Mostafa Shahin, Beena Ahmed, Julien Epps

Cognitive impairment (CI) is of growing public health concern, and early detection is vital for effective intervention. Speech has gained attention as a non-invasive and easily col…

eess.AS2025

Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction

Xiangyu Zhang, Daijiao Liu, Tianyi Xiao +5

In the speech signal, acoustic landmarks identify times when the acoustic manifestations of the linguistically motivated distinctive features are most salient. Acoustic landmarks h…

eess.AS2024

Rethinking Mamba in Speech Processing by Self-Supervised Models

Xiangyu Zhang, Jianbo Ma, Mostafa Shahin +2

The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech…