4 papers · 1 filter
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
Mohan Shi, Xiong Xiao, Ruchao Fan +2
Joint automatic speech recognition (ASR) and speaker diarization aim to answer the question "who spoke what" in multi-speaker scenarios. In this paper, we present an end-to-end spe…
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
Shaoshi Ling, Gang Liu, Guoli Ye +1
Speech summarization is a critical component of spoken content understanding, particularly in the era of rapidly growing spoken and audiovisual data. Recent advances in multi-modal…
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
Hao Yen, Shaoshi Ling, Guoli Ye
We propose a novel approach to end-to-end automatic speech recognition (ASR) to achieve efficient speech in-context learning (SICL) for (i) long-form speech decoding, (ii) test-tim…
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
Shaoshi Ling, Guoli Ye, Rui Zhao +1
The attention-based encoder-decoder (AED) speech recognition model has been widely successful in recent years. However, the joint optimization of acoustic model and language model…