activity
20182026
most citedWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

1.9k citations · 2k across the 50 of their papers we have counts for

collaborators
Showing 2021Show all

16 papers · 1 filter

eess.AS2021

Separating Long-Form Speech with Group-Wise Permutation Invariant Training

Wangyou Zhang, Zhuo Chen, Naoyuki Kanda +8

Multi-talker conversational speech processing has drawn many interests for various applications such as meeting transcription. Speech separation is often required to handle overlap…

cs.SD2021★ 1 cited

Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction

Heming Wang, Yao Qian, Xiaofei Wang +6

Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a…

eess.AS2021

Continuous Speech Separation with Recurrent Selective Attention Network

Yixuan Zhang, Zhuo Chen, Jian Wu +4

While permutation invariant training (PIT) based continuous speech separation (CSS) significantly improves the conversation transcription accuracy, it often suffers from speech lea…

eess.AS2021★ 1 cited

VarArray: Array-Geometry-Agnostic Continuous Speech Separation

Takuya Yoshioka, Xiaofei Wang, Dongmei Wang +4

Continuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper propose…

eess.AS2021

One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement

Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka +3

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect…

eess.AS2021★ 1 cited

Personalized Speech Enhancement: New Models and Comprehensive Evaluation

Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang +3

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and…