11 papers · 1 filter
Speaker Role and Language Diarization for Analyzing Multilingual Interviews for Language Proficiency of Older Adults
Anfeng Xu, Tiantian Feng, Kevin Huang +8
Automatic language proficiency assessment in the context of multilingual interview-based settings remains underexplored. In this work, we develop Whisper-based speaker-role and lan…
Speech Entrainment in Multi-Party Conversations with a Digital Agent
Nicholas Mehlman, Kaitlin Zareno, Kleanthis Avramidis +2
It has been widely observed that individuals engaged in conversation tend to adapt their speaking style to more closely match the other interlocutors. However, most prior work has…
Speech Codec Probing from Semantic and Phonetic Perspectives
Xuan Shi, Chang Zeng, Tiantian Feng +3
Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. Speech tokenizers are expected to preserve both semantic and acoustic i…
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
Simon Pistrosch, Kleanthis Avramidis, Zhao Ren +7
The expression of affect is integral to spoken communication, yet, its link to underlying articulatory execution remains unclear. Measures of articulatory muscle activity such as E…
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
Anfeng Xu, Tiantian Feng, Somer Bishop +2
Accurate transcription and speaker role diarization of child-adult spoken interactions are crucial for developmental and clinical research. However, manual annotation is time-consu…
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Helin Wang, Jiarui Hai, Dading Chong +11
Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to…