Autovocoder: Fast Waveform Generation from a Learned Speech Representation using Differentiable Digital Signal Processing
arXiv:2211.06989 · doi:10.1109/ICASSP49357.2023.10095729
Abstract
Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation. A mel-spectrogram is extracted from the waveform by a simple, fast DSP operation, but generating a high-quality waveform from a mel-spectrogram requires computationally expensive machine learning: a neural vocoder. Our proposed ``autovocoder'' reverses this arrangement. We use machine learning to obtain a representation that replaces the mel-spectrogram, and that can be inverted back to a waveform using simple, fast operations including a differentiable implementation of the inverse STFT. The autovocoder generates a waveform 5 times faster than the DSP-based Griffin-Lim algorithm, and 14 times faster than the neural vocoder HiFi-GAN. We provide perceptual listening test results to confirm that the speech is of comparable quality to HiFi-GAN in the copy synthesis task.
Accepted to the 2023 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2023)
References in corpus (7)
- WaveNet: A Generative Model for Raw Audio
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- Parallel WaveNet: Fast High-Fidelity Speech Synthesis
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
- VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
- WavThruVec: Latent speech representation as intermediate features for neural speech synthesis
- Puffin: pitch-synchronous neural waveform generation for fullband speech on modest devices