184 citations · 203 across the 6 of their papers we have counts for
4 papers · 1 filter
WaveFit: An Iterative and Non-autoregressive Neural Vocoder based on Fixed-Point Iteration
Yuma Koizumi, Kohei Yatabe, Heiga Zen +1
Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characteriz…
DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement
Yuma Koizumi, Shigeki Karita, Scott Wisdom +4
Single-channel speech enhancement (SE) is an important task in speech processing. A widely used framework combines an analysis/synthesis filterbank with a mask prediction network,…
From Audio to Semantics: Approaches to end-to-end spoken language understanding
Parisa Haghani, Arun Narayanan, Michiel Bacchiani +6
Conventional spoken language understanding systems consist of two main components: an automatic speech recognition module that converts audio to a transcript, and a natural languag…
Multi-Dialect Speech Recognition With A Single Sequence-To-Sequence Model
Bo Li, Tara N. Sainath, Khe Chai Sim +6
Sequence-to-sequence models provide a simple and elegant solution for building speech recognition systems by folding separate components of a typical system, namely acoustic (AM),…