4 papers
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
Keqi Deng, Shaoshi Ling, Ruchao Fan +1
Long-context inference in large language models (LLMs) is bottlenecked by the linear growth of the self-attention key-value (KV) cache. Top-k sparse attention alleviates this by lo…
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
Mohan Shi, Xiong Xiao, Ruchao Fan +2
Joint automatic speech recognition (ASR) and speaker diarization aim to answer the question "who spoke what" in multi-speaker scenarios. In this paper, we present an end-to-end spe…
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
Shaoshi Ling, Gang Liu, Guoli Ye +1
Speech summarization is a critical component of spoken content understanding, particularly in the era of rapidly growing spoken and audiovisual data. Recent advances in multi-modal…
Customizing Speech Recognition Model with Large Language Model Feedback
Shaoshi Ling, Guoli Ye
Automatic speech recognition (ASR) systems have achieved strong performance on general transcription tasks. However, they continue to struggle with recognizing rare named entities…