10 papers
Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis
Jinghao Chen, Mostafa Shahin, Beena Ahmed
Automatic mispronunciation detection and diagnosis (MDD) plays a crucial role in L2 Mandarin pronunciation learning. While end-to-end (E2E) based MDD methods have substantially imp…
The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models
Xiangyu Zhang, Yuxin Li, Haoyang Zhang +5
The pursuit of a "unified" discrete token for both speech understanding and generation has led the Speech Language Model (SLM) community to heavily rely on Word Error Rate (WER) --…
Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization
Xiangyu Zhang, Benjamin John Southwell, Siqi Pan +3
Audio tokenization has emerged as a critical component in end-to-end audio language models, enabling efficient discrete representation learning for both audio understanding and gen…
Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
Xiangyu Zhang, Fuming Fang, Peng Gao +3
Large Language Models (LLMs) have demonstrated remarkable success across diverse fields, establishing a powerful paradigm for complex information processing. This has inspired the…
System X: A Mobile Voice-Based AI System for EMR Generation and Clinical Decision Support in Low-Resource Maternal Healthcare
Maryam Mustafa, Umme Ammara, Amna Shahnawaz +5
We present the design, implementation, and in-situ deployment of a smartphone-based voice-enabled AI system for generating electronic medical records (EMRs) and clinical risk alert…
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
Mostafa Shahin, Beena Ahmed, Julien Epps
Cognitive impairment (CI) is of growing public health concern, and early detection is vital for effective intervention. Speech has gained attention as a non-invasive and easily col…