collaborators

8 papers

cs.SD2025

MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages

Hardik B. Sailor, Aw Ai Ti, Chen Fang Yih Nancy +26

We present MERaLiON-SER, a robust speech emotion recognition model designed for English and Southeast Asian languages. The model is trained using a hybrid objective combining weigh…

cs.CL2025

Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models

Qiongqiong Wang, Hardik B. Sailor, Jeremy H. M. Wong +6

Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextu…

cs.SD2025

ASTAR-NTU solution to AudioMOS Challenge 2025 Track1

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +3

Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically pr…

cs.SD2025

A correlation-permutation approach for speech-music encoders model merging

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong +3

Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, dire…

cs.CL2025

Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs

Wenyu Zhang, Yingxu He, Geyu Lin +9

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues s…

cs.SD2025

Distilling a speech and music encoder with task arithmetic

Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei +4

Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding…