Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Kai Li, Wenze Ren, Junjie Li +11
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols c…
eess.AS2026
Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE
Peijun Yang, Zhan Jin, Xiaoyi Qin +4
Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant v…
eess.AS2025
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
Cancan Li, Fei Su, Juan Liu +4
Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal rest…