2 citations · 2 across the 2 of their papers we have counts for
6 papers · 1 filter
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
Qiongqiong Wang, Hardik Bhupendra Sailor, Tianchi Liu +7
Recent speech-LLMs have shown impressive performance in tasks like transcription and translation, yet they remain limited in understanding the paralinguistic aspects of speech cruc…
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
Qiongqiong Wang, Hardik B. Sailor, Jeremy H. M. Wong +6
Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextu…
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
Wenyu Zhang, Yingxu He, Geyu Lin +9
Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues s…
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
Yiming Gao, Bin Wang, Chengwei Wei +2
Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after al…
MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
Yingxu He, Zhuohan Liu, Shuo Sun +5
We introduce MERaLiON-AudioLLM (Multimodal Empathetic Reasoning and Learning in One Network), the first speech-text model tailored for Singapore's multilingual and multicultural la…
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
Bin Wang, Xunlong Zou, Shuo Sun +6
Singlish, a Creole language rooted in English, is a key focus in linguistic research within multilingual and multicultural contexts. However, its spoken form remains underexplored,…