105 citations · 247 across the 22 of their papers we have counts for
29 papers
WaveFit: An Iterative and Non-autoregressive Neural Vocoder based on Fixed-Point Iteration
Yuma Koizumi, Kohei Yatabe, Heiga Zen +1
Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characteriz…
Mask scalar prediction for improving robust automatic speech recognition
Arun Narayanan, James Walker, Sankaran Panchapagesan +2
Using neural network based acoustic frontends for improving robustness of streaming automatic speech recognition (ASR) systems is challenging because of the causality constraints a…
DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement
Yuma Koizumi, Shigeki Karita, Scott Wisdom +4
Single-channel speech enhancement (SE) is an important task in speech processing. A widely used framework combines an analysis/synthesis filterbank with a mask prediction network,…
Description and Discussion on DCASE 2021 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring under Domain Shifted Conditions
Yohei Kawaguchi, Keisuke Imoto, Yuma Koizumi +6
We present the task description and discussion on the results of the DCASE 2021 Challenge Task 2. In 2020, we organized an unsupervised anomalous sound detection (ASD) task, identi…
Sampling-Frequency-Independent Audio Source Separation Using Convolution Layer Based on Impulse Invariant Method
Koichi Saito, Tomohiko Nakamura, Kohei Yatabe +2
Audio source separation is often used as preprocessing of various applications, and one of its ultimate goals is to construct a single versatile model capable of dealing with the v…
Noisy-target Training: A Training Strategy for DNN-based Speech Enhancement without Clean Speech
Takuya Fujimura, Yuma Koizumi, Kohei Yatabe +1
Deep neural network (DNN)-based speech enhancement ordinarily requires clean speech signals as the training target. However, collecting clean signals is very costly because they mu…