3 papers
eess.AS2024
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
Keyu An, Zerui Li, Zhifu Gao +1
Attention-based encoder-decoder, e.g. transformer and its variants, generates the output sequence in an autoregressive (AR) manner. Despite its superior performance, AR model is co…
cs.SD2024
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Keyu An, Qian Chen, Chong Deng +30
This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative mo…
cs.SD2024
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Zhihao Du, Jiaming Wang, Qian Chen +12
Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for a…