11 citations · 11 across the 3 of their papers we have counts for
6 papers
Leveraging Language Information for Target Language Extraction
Mehmet Sinan Yıldırım, Ruijie Tao, Wupeng Wang +2
Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory sy…
Fun-ASR Technical Report
Keyu An, Yanni Chen, Zhigao Chen +35
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep in…
Online Audio-Visual Autoregressive Speaker Extraction
Zexu Pan, Wupeng Wang, Shengkui Zhao +4
This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less ex…
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
Wupeng Wang, Zexu Pan, Xinke Li +2
Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a pro…
Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation
Wupeng Wang, Zexu Pan, Jingru Lin +2
Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences perf…
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
Wupeng Wang, Zexu Pan, Xinke Li +2
Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of ta…