11 citations · 11 across the 4 of their papers we have counts for
4 papers
Online Audio-Visual Autoregressive Speaker Extraction
Zexu Pan, Wupeng Wang, Shengkui Zhao +4
This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less ex…
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
Wupeng Wang, Zexu Pan, Xinke Li +2
Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a pro…
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
Wupeng Wang, Zexu Pan, Xinke Li +2
Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of ta…
Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech
Jingru Lin, Meng Ge, Wupeng Wang +2
Self-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-label…