Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE
Peijun Yang, Zhan Jin, Xiaoyi Qin +4
Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant v…
eess.AS2025
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
Cancan Li, Fei Su, Juan Liu +4
Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal rest…
eess.AS2024
The Database and Benchmark for the Source Speaker Tracing Challenge 2024
Ze Li, Yuke Lin, Tian Yao +6
Voice conversion (VC) systems can transform audio to mimic another speaker's voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker…