Latent-Domain Predictive Neural Speech Coding
arXiv:2207.08363 · doi:10.1109/TASLP.2023.3277693
Abstract
Neural audio/speech coding has recently demonstrated its capability to deliver high quality at much lower bitrates than traditional methods. However, existing neural audio/speech codecs employ either acoustic features or learned blind features with a convolutional neural network for encoding, by which there are still temporal redundancies within encoded features. This paper introduces latent-domain predictive coding into the VQ-VAE framework to fully remove such redundancies and proposes the TF-Codec for low-latency neural speech coding in an end-to-end manner. Specifically, the extracted features are encoded conditioned on a prediction from past quantized latent frames so that temporal correlations are further removed. Moreover, we introduce a learnable compression on the time-frequency input to adaptively adjust the attention paid to main frequencies and details at different bitrates. A differentiable vector quantization scheme based on distance-to-soft mapping and Gumbel-Softmax is proposed to better model the latent distributions with rate constraint. Subjective results on multilingual speech datasets show that, with low latency, the proposed TF-Codec at 1 kbps achieves significantly better quality than Opus at 9 kbps, and TF-Codec at 3 kbps outperforms both EVS at 9.6 kbps and Opus at 12 kbps. Numerous studies are conducted to demonstrate the effectiveness of these techniques. Code and models are available at https://github.com/microsoft/TF-Codec.
Accepted by IEEE/ACM TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING (TASLP). Code and models are available at https://github.com/microsoft/TF-Codec
References in corpus (14)
- Adam: A Method for Stochastic Optimization
- WaveNet: A Generative Model for Raw Audio
- Neural Discrete Representation Learning
- Empirical Evaluation of Rectified Activations in Convolutional Network
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- High Fidelity Neural Audio Compression
- Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder
- Deep Contextual Video Compression
- Scalable and Efficient Neural Speech Coding: A Hybrid Design
Cited by in corpus (5)
- Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
- SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
- Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
- Personalized Neural Speech Codec
- Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision