4 papers
Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE
Peijun Yang, Zhan Jin, Xiaoyi Qin +4
Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant v…
Multi-View Based Audio Visual Target Speaker Extraction
Peijun Yang, Zhan Jin, Juan Liu +1
Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE me…
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
Zhan Jin, Bang Zeng, Peijun Yang +5
Audio-Visual Target Speaker Extraction (AVTSE) is crucial for cocktail party scenarios. Leveraging multiple cues --such as utterance-level speaker embeddings or steady face images,…
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
Jiarong Du, Zhan Jin, Peijun Yang +4
Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often…