2 papers
eess.AS2025
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation
Chunyu Sun, Bingyu Liu, Zhichao Cui +5
Embedding-based retrieval models have made significant strides in retrieval-augmented generation (RAG) techniques for text and multimodal large language models (LLMs) applications.…
cs.SD2025
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
Shu Wu, Anbin Qi, Yanzhang Xie +1
Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled spee…