45 citations · 112 across the 22 of their papers we have counts for
29 papers
Streaming Target-Speaker ASR with Neural Transducer
Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai +2
Although recent advances in deep learning technology have boosted automatic speech recognition (ASR) performance in the single-talker case, it remains difficult to recognize multi-…
Mask-based Neural Beamforming for Moving Speakers with Self-Attention-based Tracking
Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani +1
Beamforming is a powerful tool designed to enhance speech signals from the direction of a target source. Computing the beamforming filter requires estimating spatial covariance mat…
Tight integration of neural- and clustering-based diarization through deep unfolding of infinite Gaussian mixture model
Keisuke Kinoshita, Marc Delcroix, Tomoharu Iwata
Speaker diarization has been investigated extensively as an important central task for meeting analysis. Recent trend shows that integration of end-to-end neural (EEND)-and cluster…
Speeding Up Permutation Invariant Training for Source Separation
Thilo von Neumann, Christoph Boeddeker, Keisuke Kinoshita +2
Permutation invariant training (PIT) is a widely used training criterion for neural network-based source separation, used for both utterance-level separation with utterance-level P…
Graph-PIT: Generalized permutation invariant training for continuous separation of arbitrary numbers of speakers
Thilo von Neumann, Keisuke Kinoshita, Christoph Boeddeker +2
Automatic transcription of meetings requires handling of overlapped speech, which calls for continuous speech separation (CSS) systems. The uPIT criterion was proposed for utteranc…
Few-shot learning of new sound classes for target sound extraction
Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai +2
Target sound extraction consists of extracting the sound of a target acoustic event (AE) class from a mixture of AE sounds. It can be realized using a neural network that extracts…