2 citations · 10 across the 15 of their papers we have counts for
18 papers
Speech separation with large-scale self-supervised learning
Zhuo Chen, Naoyuki Kanda, Jian Wu +6
Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…
Simulating realistic speech overlaps improves multi-talker ASR
Muqiao Yang, Naoyuki Kanda, Xiaofei Wang +5
Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the di…
Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition
Zili Huang, Zhuo Chen, Naoyuki Kanda +6
Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…
VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition
Naoyuki Kanda, Jian Wu, Xiaofei Wang +3
This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geom…
Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization
Dongmei Wang, Xiong Xiao, Naoyuki Kanda +2
This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model's drawback of b…
The Microsoft System for VoxCeleb Speaker Recognition Challenge 2022
Gang Liu, Tianyan Zhou, Yong Zhao +4
In this report, we describe our submitted system for track 2 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). We fuse a variety of good-performing models ranging fro…