collaborators

12 papers

eess.AS2026

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE

Peijun Yang, Zhan Jin, Xiaoyi Qin +4

Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant v…

eess.AS2026

ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection

Yucong Zhang, Juan Liu, Ming Li

Machine anomalous sound detection (ASD) requires robust audio representations capable of capturing subtle deviations in machine sounds under limited supervision. Existing pre-train…

eess.AS2026

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models

Li Li, Ming Cheng, Weixin Zhu +3

Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but al…

eess.AS2026

Multi-View Based Audio Visual Target Speaker Extraction

Peijun Yang, Zhan Jin, Juan Liu +1

Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE me…

eess.AS2026

Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion

Zhan Jin, Bang Zeng, Peijun Yang +5

Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images,…

eess.AS2026

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion

Dong Liu, Juan Liu, Wei Ju +2

Whispered speech lacks vocal-fold excitation, making intelligible conversion challenging. We propose WhisperVC, a three-stage framework for low-resource whisper-to-normal (W2N) con…