6 papers
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
Weiqiao Shan, Yuang Li, Yuhao Zhang +9
Connecting audio encoders with large language models (LLMs) allows the LLM to perform various audio understanding tasks, such as automatic speech recognition (ASR) and audio captio…
Why Not Transform Chat Large Language Models to Non-English?
Xiang Geng, Ming Zhu, Jiahuan Li +14
The scarcity of non-English data limits the development of non-English large language models (LLMs). Transforming English-centric LLMs to non-English has been identified as an effe…
DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation
Xinglin Lyu, Wei Tang, Yuang Li +7
Document-level context is crucial for handling discourse challenges in text-to-text document-level machine translation (MT). Despite the increased discourse challenges introduced b…
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
Weiqiao Shan, Yuhao Zhang, Yuchen Han +7
Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech…
Investigating Numerical Translation with Large Language Models
Wei Tang, Jiawei Yu, Yuang Li +5
The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made…
"I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities
Jiawei Yu, Xiang Geng, Yuang Li +8
Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, ann…