12 citations · 45 across the 15 of their papers we have counts for
13 papers
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
Yoshiki Masuyama, Natsuki Ueno, Nobutaka Ono
Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propos…
Online Phase Reconstruction via DNN-based Phase Differences Estimation
Yoshiki Masuyama, Kohei Yatabe, Kento Nagatomo +1
This paper presents a two-stage online phase reconstruction framework using causal deep neural networks (DNNs). Phase reconstruction is a task of recovering phase of the short-time…
End-to-End Integration of Speech Recognition, Dereverberation, Beamforming, and Self-Supervised Learning Representation
Yoshiki Masuyama, Xuankai Chang, Samuele Cornell +2
Self-supervised learning representation (SSLR) has demonstrated its significant effectiveness in automatic speech recognition (ASR), mainly with clean speech. Recent work pointed o…
Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
Yoshiki Masuyama, Yoshiaki Bando, Kohei Yatabe +3
Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with…
Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention
Yuma Koizumi, Kohei Yatabe, Marc Delcroix +2
This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly fro…
Phase reconstruction based on recurrent phase unwrapping with deep neural networks
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi +2
Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio s…