36 papers
Speaker Role and Language Diarization for Analyzing Multilingual Interviews for Language Proficiency of Older Adults
Anfeng Xu, Tiantian Feng, Kevin Huang +8
Automatic language proficiency assessment in the context of multilingual interview-based settings remains underexplored. In this work, we develop Whisper-based speaker-role and lan…
Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
Shih-Heng Wang, Tiantian Feng, Aditya Kommineni +4
Neural Audio Codecs (NACs) are widely adopted in modern speech systems, yet how they encode linguistic and paralinguistic information remains unclear. Improving the interpretabilit…
Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs
Efthymios Tsaprazlis, Tiantian Feng, Anil Ramakrishna +3
Existing visual privacy benchmarks largely treat privacy as a binary property, labeling images as private or non-private based on visible sensitive content. We argue that privacy i…
Speech Codec Probing from Semantic and Phonetic Perspectives
Xuan Shi, Chang Zeng, Tiantian Feng +3
Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. Speech tokenizers are expected to preserve both semantic and acoustic i…
An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
Jihwan Lee, Parsa Razmara, Kevin Huang +16
Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the acoustic speech signal is the most accessi…
RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection
Wenze Ren, Ke-Han Lu, Kai-Wei Chang +8
Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) e…