9 papers · 1 filter
MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages
Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen +18
We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes M…
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham +3
Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, w…
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
Yiming Gao, Bin Wang, Chengwei Wei +2
Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after al…
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
Wenyu Zhang, Yingxu He, Geyu Lin +9
Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues s…
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
Qiongqiong Wang, Hardik Bhupendra Sailor, Tianchi Liu +7
Recent speech-LLMs have shown impressive performance in tasks like transcription and translation, yet they remain limited in understanding the paralinguistic aspects of speech cruc…
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
Qiongqiong Wang, Hardik B. Sailor, Jeremy H. M. Wong +6
Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextu…