5 citations · 8 across the 7 of their papers we have counts for
7 papers
Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization
Dongmei Wang, Xiong Xiao, Naoyuki Kanda +2
This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model's drawback of b…
Leveraging Real Conversational Data for Multi-Channel Continuous Speech Separation
Xiaofei Wang, Dongmei Wang, Naoyuki Kanda +2
Existing multi-channel continuous speech separation (CSS) models are heavily dependent on supervised data - either simulated data which causes data mismatch between the training an…
PickNet: Real-Time Channel Selection for Ad Hoc Microphone Arrays
Takuya Yoshioka, Xiaofei Wang, Dongmei Wang
This paper proposes PickNet, a neural network model for real-time channel selection for an ad hoc microphone array consisting of multiple recording devices like cell phones. Assumi…
VarArray: Array-Geometry-Agnostic Continuous Speech Separation
Takuya Yoshioka, Xiaofei Wang, Dongmei Wang +4
Continuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper propose…
All-neural beamformer for continuous speech separation
Zhuohuang Zhang, Takuya Yoshioka, Naoyuki Kanda +4
Continuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unkn…
Continuous Speech Separation with Ad Hoc Microphone Arrays
Dongmei Wang, Takuya Yoshioka, Zhuo Chen +3
Speech separation has been shown effective for multi-talker speech recognition. Under the ad hoc microphone array setup where the array consists of spatially distributed asynchrono…