4 papers
Spectrogram features for audio and speech analysis
Ian McLoughlin, Lam Pham, Yan Song +7
Spectrogram-based representations have grown to dominate the feature space for deep learning audio analysis systems, and are often adopted for speech analysis also. Initially, the…
Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries
Tushar Pranav, Eshan Pandey, Austria Lyka Diane Bala +3
Vision-Language Models (VLMs) excel in multimodal tasks but often exhibit Western-centric biases, limiting their effectiveness in culturally diverse regions like Southeast Asia (SE…
Automated evaluation of children's speech fluency for low-resource languages
Bowen Zhang, Nur Afiqah Abdul Latiff, Justin Kan +4
Assessment of children's speaking fluency in education is well researched for majority languages, but remains highly challenging for low resource languages. This paper proposes a s…
Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
Xiaoxiao Miao, Yuxiang Zhang, Xin Wang +3
A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores h…