activity
20162026
most citedWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

1.9k citations · 1.9k across the 14 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

eess.AS2021

Separating Long-Form Speech with Group-Wise Permutation Invariant Training

Wangyou Zhang, Zhuo Chen, Naoyuki Kanda +8

Multi-talker conversational speech processing has drawn many interests for various applications such as meeting transcription. Speech separation is often required to handle overlap…

cs.CL2021★ 1.9k cited

WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

Sanyuan Chen, Chengyi Wang, Zhengyang Chen +15

Self-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks. As speech signal contai…

eess.AS2021★ 3 cited

Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR

Naoyuki Kanda, Xiong Xiao, Yashesh Gaur +4

This paper presents Transcribe-to-Diarize, a new approach for neural speaker diarization that uses an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR). The…

eess.AS2021★ 1 cited

A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio

Naoyuki Kanda, Xiong Xiao, Jian Wu +6

Speaker-attributed automatic speech recognition (SA-ASR) is a task to recognize "who spoke what" from multi-talker recordings. An SA-ASR system usually consists of multiple modules…

eess.AS2021★ 12 cited

Speaker attribution with voice profiles by graph-based semi-supervised learning

Jixuan Wang, Xiong Xiao, Jian Wu +3

Speaker attribution is required in many real-world applications, such as meeting transcription, where speaker identity is assigned to each utterance according to speaker voice prof…