12 papers
Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE
Peijun Yang, Zhan Jin, Xiaoyi Qin +4
Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant v…
ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection
Yucong Zhang, Juan Liu, Ming Li
Machine anomalous sound detection (ASD) requires robust audio representations capable of capturing subtle deviations in machine sounds under limited supervision. Existing pre-train…
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
Li Li, Ming Cheng, Weixin Zhu +3
Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but al…
Multi-View Based Audio Visual Target Speaker Extraction
Peijun Yang, Zhan Jin, Juan Liu +1
Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE me…
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
Zhan Jin, Bang Zeng, Peijun Yang +5
Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images,…
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
Dong Liu, Juan Liu, Wei Ju +2
Whispered speech lacks vocal-fold excitation, making intelligible conversion challenging. We propose WhisperVC, a three-stage framework for low-resource whisper-to-normal (W2N) con…