6 citations · 12 across the 7 of their papers we have counts for
6 papers · 1 filter
Label-Synchronous Speech-to-Text Alignment for ASR Using Forward and Backward Transformers
Yusuke Kida, Tatsuya Komatsu, Masahito Togami
This paper proposes a novel label-synchronous speech-to-text alignment technique for automatic speech recognition (ASR). The speech-to-text alignment is a problem of splitting long…
Joint Dereverberation and Separation with Iterative Source Steering
Taishi Nakashima, Robin Scheibler, Masahito Togami +1
We propose a new algorithm for joint dereverberation and blind source separation (DR-BSS). Our work builds upon the IRLMA-T framework that applies a unified filter combining dereve…
Surrogate Source Model Learning for Determined Source Separation
Robin Scheibler, Masahito Togami
We propose to learn surrogate functions of universal speech priors for determined blind speech separation. Deep speech priors are highly desirable due to their high modelling power…
Consistency-aware multi-channel speech enhancement using deep neural networks
Yoshiki Masuyama, Masahito Togami, Tatsuya Komatsu
This paper proposes a deep neural network (DNN)-based multi-channel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal.…
Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function
Masahito Togami, Yoshiki Masuyama, Tatsuya Komatsu +1
In this paper, we propose a multi-channel speech source separation with a deep neural network (DNN) which is trained under the condition that no clean signal is available. As an al…
Multi-channel Time-Varying Covariance Matrix Model for Late Reverberation Reduction
Masahito Togami
In this paper, a multi-channel time-varying covariance matrix model for late reverberation reduction is proposed. Reflecting that variance of the late reverberation is time-varying…