1 paper · 1 filter
Peijun Yang, Zhan Jin, Xiaoyi Qin +4
Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant v…