activity
20192022
most citedJointly optimal denoising, dereverberation, and source separation

45 citations · 112 across the 22 of their papers we have counts for

collaborators

29 papers

eess.AS2022

Streaming Target-Speaker ASR with Neural Transducer

Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai +2

Although recent advances in deep learning technology have boosted automatic speech recognition (ASR) performance in the single-talker case, it remains difficult to recognize multi-…

eess.AS20221 cited

Mask-based Neural Beamforming for Moving Speakers with Self-Attention-based Tracking

Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani +1

Beamforming is a powerful tool designed to enhance speech signals from the direction of a target source. Computing the beamforming filter requires estimating spatial covariance mat…

eess.AS2022

Tight integration of neural- and clustering-based diarization through deep unfolding of infinite Gaussian mixture model

Keisuke Kinoshita, Marc Delcroix, Tomoharu Iwata

Speaker diarization has been investigated extensively as an important central task for meeting analysis. Recent trend shows that integration of end-to-end neural (EEND)-and cluster…

eess.AS20216 cited

Speeding Up Permutation Invariant Training for Source Separation

Thilo von Neumann, Christoph Boeddeker, Keisuke Kinoshita +2

Permutation invariant training (PIT) is a widely used training criterion for neural network-based source separation, used for both utterance-level separation with utterance-level P…

eess.AS202120 cited

Graph-PIT: Generalized permutation invariant training for continuous separation of arbitrary numbers of speakers

Thilo von Neumann, Keisuke Kinoshita, Christoph Boeddeker +2

Automatic transcription of meetings requires handling of overlapped speech, which calls for continuous speech separation (CSS) systems. The uPIT criterion was proposed for utteranc…

eess.AS2021

Few-shot learning of new sound classes for target sound extraction

Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai +2

Target sound extraction consists of extracting the sound of a target acoustic event (AE) class from a mixture of AE sounds. It can be realized using a neural network that extracts…