activity
20172023
most citedWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

1.9k citations · 2.7k across the 81 of their papers we have counts for

collaborators
Showing 2022 · eess.ASShow all

11 papers · 2 filters

eess.AS2022

Speech separation with large-scale self-supervised learning

Zhuo Chen, Naoyuki Kanda, Jian Wu +6

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…

eess.AS2022★ 1 cited

Simulating realistic speech overlaps improves multi-talker ASR

Muqiao Yang, Naoyuki Kanda, Xiaofei Wang +5

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the di…

eess.AS2022

Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition

Zili Huang, Zhuo Chen, Naoyuki Kanda +6

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…

eess.AS2022★ 36 cited

VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning

Qiushi Zhu, Long Zhou, Ziqiang Zhang +7

Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, t…

eess.AS2022★ 1 cited

Acoustic-aware Non-autoregressive Spell Correction with Mask Sample Decoding

Ruchao Fan, Guoli Ye, Yashesh Gaur +1

Masked language model (MLM) has been widely used for understanding tasks, e.g. BERT. Recently, MLM has also been used for generation tasks. The most popular one in speech is using…

eess.AS2022★ 2 cited

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

Naoyuki Kanda, Jian Wu, Xiaofei Wang +3

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geom…