Variable-rate discrete representation learning
arXiv:2103.06089
Abstract
Semantically meaningful information content in perceptual signals is usually unevenly distributed. In speech signals for example, there are often many silences, and the speed of pronunciation can vary considerably. In this work, we propose slow autoencoders (SlowAEs) for unsupervised learning of high-level variable-rate discrete representations of sequences, and apply them to speech. We show that the resulting event-based representations automatically grow or shrink depending on the density of salient information in the input signals, while still allowing for faithful signal reconstruction. We develop run-length Transformers (RLTs) for event-based representation modelling and use them to construct language models in the speech domain, which are able to generate grammatical and semantically coherent utterances and continuations.
26 pages, 15 figures, samples can be found at https://vdrl.github.io/
References in corpus (28)
- Auto-Encoding Variational Bayes
- Categorical Reparameterization with Gumbel-Softmax
- Language Models are Few-Shot Learners
- Neural Discrete Representation Learning
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Libri-Light: A Benchmark for ASR with Limited or No Supervision
- Visual Transformers: Token-based Image Representation and Processing for Computer Vision
- Adaptive Computation Time for Recurrent Neural Networks
- Variational Lossy Autoencoder
- Multi-Object Representation Learning with Iterative Variational Inference
- MONet: Unsupervised Scene Decomposition and Representation
- VideoBERT: A Joint Model for Video and Language Representation Learning
- The challenge of realistic music generation: modelling raw audio at scale
- Jukebox: A Generative Model for Music
- SOM-VAE: Interpretable Discrete Representation Learning on Time Series
- Generating High Fidelity Images with Subscale Pixel Networks and Multidimensional Upscaling
- Learning Ordered Representations with Nested Dropout
- Representation Learning using Event-based STDP
- Discrete Event, Continuous Time RNNs
- Interpolation-Prediction Networks for Irregularly Sampled Time Series
- Hierarchical Autoregressive Image Models with Auxiliary Decoders
- DiscreTalk: Text-to-Speech as a Machine Translation Problem
- Efficient Segmentation: Learning Downsampling Near Semantic Boundaries
- Vector Quantized Contrastive Predictive Coding for Template-based Music Generation
- Generating Images with Sparse Representations
- Surprisal-Triggered Conditional Computation with Neural Networks