Learning and controlling the source-filter representation of speech with a variational autoencoder
arXiv:2204.07075 · doi:10.1016/j.specom.2023.02.005
Abstract
Understanding and controlling latent representations in deep generative models is a challenging yet important problem for analyzing, transforming and generating various types of data. In speech processing, inspiring from the anatomical mechanisms of phonation, the source-filter model considers that speech signals are produced from a few independent and physically meaningful continuous latent factors, among which the fundamental frequency and the formants are of primary importance. In this work, we start from a variational autoencoder (VAE) trained in an unsupervised manner on a large dataset of unlabeled natural speech signals, and we show that the source-filter model of speech production naturally arises as orthogonal subspaces of the VAE latent space. Using only a few seconds of labeled speech signals generated with an artificial speech synthesizer, we propose a method to identify the latent subspaces encoding and the first three formant frequencies, we show that these subspaces are orthogonal, and based on this orthogonality, we develop a method to accurately and independently control the source-filter speech factors within the latent subspaces. Without requiring additional information such as text or human-labeled data, this results in a deep generative model of speech spectrograms that is conditioned on and the formant frequencies, and which is applied to the transformation speech signals. Finally, we also propose a robust estimation method that exploits the projection of a speech signal onto the learned latent subspace associated with .
23 pages, 7 figures, companion website: https://samsad35.github.io/site-sfvae/
References in corpus (9)
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
- Controlling generative models with continuous factors of variations
- Deep Learning Based Assessment of Synthetic Speech Naturalness
- A variance modeling framework based on variational autoencoders for speech enhancement
- Disentanglement by Nonlinear ICA with General Incompressible-flow Networks (GIN)
- Speech enhancement with variational autoencoders and alpha-stable distributions
- A Sober Look at the Unsupervised Learning of Disentangled Representations and their Evaluation
- Neural Pitch-Shifting and Time-Stretching with Controllable LPCNet
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations