3 papers
cs.SD2025
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
Shu Wu, Anbin Qi, Yanzhang Xie +1
Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled spee…
eess.AS2025
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation
Chunyu Sun, Bingyu Liu, Zhichao Cui +5
Embedding-based retrieval models have made significant strides in retrieval-augmented generation (RAG) techniques for text and multimodal large language models (LLMs) applications.…
cs.CV2024
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
Anbin QI, Zhongliang Liu, Xinyong Zhou +6
In this paper, we present our solution for the Second Multimodal Emotion Recognition Challenge Track 1(MER2024-SEMI). To enhance the accuracy and generalization performance of emot…