4 papers · 1 filter
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
Wenqian Cui, Xiao-Hui Li, Daxin Tan +2
Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial modality gap. Prior work has main…
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
Jing Xu, Jiaqi Wang, Daxin Tan +1
Although Large Language Models (LLMs) excel in many tasks, their application to Speech-to-Speech Translation (S2ST) is underexplored and hindered by data scarcity. To bridge this g…
AEQ-Bench: Measuring Empathy of Omni-Modal Large Models
Xuan Luo, Lewei Yao, Libo Zhao +6
While the automatic evaluation of omni-modal large models (OLMs) is essential, assessing empathy remains a significant challenge due to its inherent affectivity. To investigate thi…
Exploring SSL Discrete Tokens for Multilingual ASR
Mingyu Cui, Daxin Tan, Yifan Yang +5
With the advancement of Self-supervised Learning (SSL) in speech-related tasks, there has been growing interest in utilizing discrete tokens generated by SSL for automatic speech r…