Neural Discrete Representation Learning
arXiv:1711.00937
Abstract
Learning useful representations without supervision remains a key challenge in machine learning. In this paper, we propose a simple yet powerful generative model that learns such discrete representations. Our model, the Vector Quantised-Variational AutoEncoder (VQ-VAE), differs from VAEs in two key ways: the encoder network outputs discrete, rather than continuous, codes; and the prior is learnt rather than static. In order to learn a discrete latent representation, we incorporate ideas from vector quantisation (VQ). Using the VQ method allows the model to circumvent issues of "posterior collapse" -- where the latents are ignored when they are paired with a powerful autoregressive decoder -- typically observed in the VAE framework. Pairing these representations with an autoregressive prior, the model can generate high quality images, videos, and speech as well as doing high quality speaker conversion and unsupervised learning of phonemes, providing further evidence of the utility of the learnt representations.
Cited by in corpus (428)
- A Survey on Visual Transformer
- On the Opportunities and Risks of Foundation Models
- Diffusion Models Beat GANs on Image Synthesis
- Diffusion Models in Vision: A Survey
- Zero-Shot Text-to-Image Generation
- Self-supervised Learning: Generative or Contrastive
- BEiT: BERT Pre-Training of Image Transformers
- Blended Diffusion for Text-driven Editing of Natural Images
- Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
- Cascaded Diffusion Models for High Fidelity Image Generation
- Model-Based Reinforcement Learning for Atari
- CogView: Mastering Text-to-Image Generation via Transformers
- Large Scale Adversarial Representation Learning
- Recent Advances in Autoencoder-Based Representation Learning
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Uncertainty Estimation Using a Single Deep Deterministic Neural Network
- Neural Image Compression for Gigapixel Histopathology Image Analysis
- Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron
- Unsupervised speech representation learning using WaveNet autoencoders
- Leveraging Frequency Analysis for Deep Fake Image Recognition
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
- Masked Autoencoders Are Scalable Vision Learners
- Dynamical Variational Autoencoders: A Comprehensive Review
- EfficientFi: Towards Large-Scale Lightweight WiFi Sensing via CSI Compression
- SpaText: Spatio-Textual Representation for Controllable Image Generation
- Scaling Laws for Autoregressive Generative Modeling
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Plan-Structured Deep Neural Network Models for Query Performance Prediction
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
- Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder
- Generative AI for Synthetic Data Across Multiple Medical Modalities: A Systematic Review of Recent Developments and Challenges
- MelNet: A Generative Model for Audio in the Frequency Domain
- Jukebox: A Generative Model for Music
- Advancements in Point Cloud Data Augmentation for Deep Learning: A Survey
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- From Variational to Deterministic Autoencoders
- Time Series Diffusion Method: A Denoising Diffusion Probabilistic Model for Vibration Signal Generation
- Face Generation and Editing with StyleGAN: A Survey
- Learning Disentangled Representations in the Imaging Domain
- Effectiveness of self-supervised pre-training for speech recognition
- CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders
- Vector-quantized Image Modeling with Improved VQGAN
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- Diagnosing and Enhancing VAE Models
- Towards a universal neural network encoder for time series
- Adaptive Latent Diffusion Model for 3D Medical Image to Image Translation: Multi-modal Magnetic Resonance Imaging Study
- Data synthesis and adversarial networks: A review and meta-analysis in cancer imaging
- A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Bayesian Layers: A Module for Neural Network Uncertainty
- AudioDec: An Open-source Streaming High-fidelity Neural Audio Codec
- LDMVFI: Video Frame Interpolation with Latent Diffusion Models
- Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
- LMQFormer: A Laplace-Prior-Guided Mask Query Transformer for Lightweight Snow Removal
- Txt2Img-MHN: Remote Sensing Image Generation from Text Using Modern Hopfield Networks
- Improving Inference for Neural Image Compression
- Lossy Image Compression with Quantized Hierarchical VAEs
- Context Disentangling and Prototype Inheriting for Robust Visual Grounding
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
- M6: A Chinese Multimodal Pretrainer
- StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks
- Unsupervised Paraphrase Generation using Pre-trained Language Models
- Explainable Time Series Anomaly Detection using Masked Latent Generative Modeling
- Artificial intelligence approaches for materials-by-design of energetic materials: state-of-the-art, challenges, and future directions
- A Universal Music Translation Network
- Robust Training of Vector Quantized Bottleneck Models
- Container: Context Aggregation Network
- Vector Quantized Diffusion Model for Text-to-Image Synthesis
- Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
- Learning from Few Samples: A Survey
- DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation
- Online Learned Continual Compression with Adaptive Quantization Modules
- LAFITE: Towards Language-Free Training for Text-to-Image Generation
- RockGPT: Reconstructing three-dimensional digital rocks from single two-dimensional slice from the perspective of video generation
- Avoiding Latent Variable Collapse With Generative Skip Models
- Audio Transformers
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- MatFuse: Controllable Material Generation with Diffusion Models
- Latent Diffusion Model for Conditional Reservoir Facies Generation
- A Tutorial on Deep Latent Variable Models of Natural Language
- Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion
- Autoregressive Quantile Networks for Generative Modeling
- Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding
- Self-Supervised VQ-VAE for One-Shot Music Style Transfer
- CCVS: Context-aware Controllable Video Synthesis
- Efficient automatic design of robots
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Feature Quantization Improves GAN Training
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design
- An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
- Scalable and Efficient Neural Speech Coding: A Hybrid Design
- Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
- Hierarchical Quantized Autoencoders
- Predicting Video with VQVAE
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
- Hierarchical Autoregressive Image Models with Auxiliary Decoders
- Latent-Domain Predictive Neural Speech Coding
- Learning Latent Plans from Play
- Self-Supervised Learning with Kernel Dependence Maximization
- Poincaré Wasserstein Autoencoder
- An Empirical Study of Training End-to-End Vision-and-Language Transformers
- Differentiable JPEG: The Devil is in the Details
- Codified audio language modeling learns useful representations for music information retrieval
- Unsupervised Discrete Sentence Representation Learning for Interpretable Neural Dialog Generation
- DiscreTalk: Text-to-Speech as a Machine Translation Problem
- Exploiting Negative Preference in Content-based Music Recommendation with Contrastive Learning
- TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial Training
- Taming Visually Guided Sound Generation
- A multimodal dynamical variational autoencoder for audiovisual speech representation learning
- MALA: Cross-Domain Dialogue Generation with Action Learning
- A Survey on Self-supervised Pre-training for Sequential Transfer Learning in Neural Networks
- A RAD approach to deep mixture models
- Generative Latent Flow
- On the Convergence of AdaBound and its Connection to SGD
- Aligned Contrastive Predictive Coding
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons
- A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural Networks
- Motif-Driven Contrastive Learning of Graph Representations
- Super-Resolving Face Image by Facial Parsing Information
- Text-Free Prosody-Aware Generative Spoken Language Modeling
- A Review of Learning with Deep Generative Models from Perspective of Graphical Modeling
- A Tutorial on VAEs: From Bayes' Rule to Lossless Compression
- VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention
- SimMIM: A Simple Framework for Masked Image Modeling
- Brain2Word: Decoding Brain Activity for Language Generation
- VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net architecture
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- Associative Compression Networks for Representation Learning
- Progressive Tandem Learning for Pattern Recognition with Deep Spiking Neural Networks
- Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements
- Prototype-based Neural Network Layers: Incorporating Vector Quantization
- Preventing Posterior Collapse with delta-VAEs
- Discrete-Valued Neural Communication
- Non-Parallel Voice Conversion with Cyclic Variational Autoencoder
- Novelty Detection Via Blurring
- Maximizing Mutual Information for Tacotron
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations
- Score-based Generative Modeling in Latent Space
- 3D Brain and Heart Volume Generative Models: A Survey
- Competitive Training of Mixtures of Independent Deep Generative Models
- VQVAE Unsupervised Unit Discovery and Multi-scale Code2Spec Inverter for Zerospeech Challenge 2019
- Tiny Transformers for Environmental Sound Classification at the Edge
- Learning to Hash with Graph Neural Networks for Recommender Systems
- Neuromorphologicaly-preserving Volumetric data encoding using VQ-VAE
- Implicit Rank-Minimizing Autoencoder
- VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector Quantization
- Vector Quantized Contrastive Predictive Coding for Template-based Music Generation
- Transformers predicting the future. Applying attention in next-frame and time series forecasting
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- Bayesian Image Reconstruction using Deep Generative Models
- Adversarial Feature Learning and Unsupervised Clustering based Speech Synthesis for Found Data with Acoustic and Textual Noise
- Discrete Representations Strengthen Vision Transformer Robustness
- Variable-rate discrete representation learning
- Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE
- Source Separation with Deep Generative Priors
- Residual Correction in Real-Time Traffic Forecasting
- Variational Autoencoders with Riemannian Brownian Motion Priors
- High Fidelity Face Manipulation with Extreme Poses and Expressions
- Latent Adversarial Debiasing: Mitigating Collider Bias in Deep Neural Networks
- Information Maximizing Visual Question Generation
- A vector quantized masked autoencoder for audiovisual speech emotion recognition
- Regularized Autoencoders via Relaxed Injective Probability Flow
- Exploring TTS without T Using Biologically/Psychologically Motivated Neural Network Modules (ZeroSpeech 2020)
- 3D-RETR: End-to-End Single and Multi-View 3D Reconstruction with Transformers
- COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
- Generating Images with Sparse Representations
- A Contrastive Learning Approach for Training Variational Autoencoder Priors
- Autoencoding sensory substitution
- Learning Discrete Structured Representations by Adversarially Maximizing Mutual Information
- FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion
- Variational Variance: Simple, Reliable, Calibrated Heteroscedastic Noise Variance Parameterization
- (Self-Attentive) Autoencoder-based Universal Language Representation for Machine Translation
- Disentangling and Learning Robust Representations with Natural Clustering
- SingSong: Generating musical accompaniments from singing
- VQ-NeRF: Neural Reflectance Decomposition and Editing with Vector Quantization
- Soft then Hard: Rethinking the Quantization in Neural Image Compression
- Predictive Sampling with Forecasting Autoregressive Models
- Weakly Supervised Disentangled Representation for Goal-conditioned Reinforcement Learning
- Video-Text Pre-training with Learned Regions
- Video-Guided Curriculum Learning for Spoken Video Grounding
- Locally Masked Convolution for Autoregressive Models
- Gaussian mixture models with Wasserstein distance
- Variational Autoencoders for Jet Simulation
- On Perceptual Lossy Compression: The Cost of Perceptual Reconstruction and An Optimal Training Framework
- Vector-Quantized Autoregressive Predictive Coding
- NoiseVC: Towards High Quality Zero-Shot Voice Conversion
- Latent Normalizing Flows for Many-to-Many Cross-Domain Mappings
- OCTOPUS: Overcoming Performance andPrivatization Bottlenecks in Distributed Learning
- Continuous Mixtures of Tractable Probabilistic Models
- MFCCGAN: A Novel MFCC-Based Speech Synthesizer Using Adversarial Learning
- Anytime Sampling for Autoregressive Models via Ordered Autoencoding
- Discovering Dialog Structure Graph for Open-Domain Dialog Generation
- DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
- Variational Inference In Pachinko Allocation Machines
- Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes
- Communication-Computation Trade-Off in Resource-Constrained Edge Inference
- Pythae: Unifying Generative Autoencoders in Python -- A Benchmarking Use Case
- Generative Models for Improved Naturalness, Intelligibility, and Voicing of Whispered Speech
- Attention Approximates Sparse Distributed Memory
- CMTS: Conditional Multiple Trajectory Synthesizer for Generating Safety-critical Driving Scenarios
- Transfer Learning with Jukebox for Music Source Separation
- tvGP-VAE: Tensor-variate Gaussian Process Prior Variational Autoencoder
- Variational Autoencoder with Implicit Optimal Priors
- Unsupervised Paraphrasing by Simulated Annealing
- Vector-Quantized Timbre Representation
- Manifolds for Unsupervised Visual Anomaly Detection
- Variational Inference for Data-Efficient Model Learning in POMDPs
- Deep Auto-encoder with Neural Response
- A survey on Variational Autoencoders from a GreenAI perspective
- An Exploratory Study on Perceptual Spaces of the Singing Voice
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- Unsupervised Subword Modeling Using Autoregressive Pretraining and Cross-Lingual Phone-Aware Modeling
- Latent Programmer: Discrete Latent Codes for Program Synthesis
- StarGAN-VC+ASR: StarGAN-based Non-Parallel Voice Conversion Regularized by Automatic Speech Recognition
- Spatial PixelCNN: Generating Images from Patches
- Max-Affine Spline Insights into Deep Generative Networks
- Failure Modes of Variational Autoencoders and Their Effects on Downstream Tasks
- Spectrogram Inpainting for Interactive Generation of Instrument Sounds
- A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery
- FastVC: Fast Voice Conversion with non-parallel data
- Target-Embedding Autoencoders for Supervised Representation Learning
- AGAIN-VC: A One-shot Voice Conversion using Activation Guidance and Adaptive Instance Normalization
- fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
- Modulating human brain responses via optimal natural image selection and synthetic image generation
- Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation
- P-Hologen: An End-to-End Generative Framework for Phase-Only Holograms
- Self-supervised Tumor Segmentation through Layer Decomposition
- Direct Evolutionary Optimization of Variational Autoencoders With Binary Latents
- Sliced Iterative Normalizing Flows
- Learning Latent Space Energy-Based Prior Model
- Unsupervised Audiovisual Synthesis via Exemplar Autoencoders
- Predictive Coding for Boosting Deep Reinforcement Learning with Sparse Rewards
- Select and Attend: Towards Controllable Content Selection in Text Generation
- Entropy Minimization In Emergent Languages
- Extractive Summary as Discrete Latent Variables
- Variational Open-Domain Question Answering
- Continual Learning: Tackling Catastrophic Forgetting in Deep Neural Networks with Replay Processes
- M3D-GAN: Multi-Modal Multi-Domain Translation with Universal Attention
- Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
- Deep Retrieval: Learning A Retrievable Structure for Large-Scale Recommendations
- Unpaired Image-to-Image Translation via Latent Energy Transport
- SketchBetween: Video-to-Video Synthesis for Sprite Animation via Sketches
- Exponential Tilting of Generative Models: Improving Sample Quality by Training and Sampling from Latent Energy
- Improved Transformer for High-Resolution GANs
- Dueling Decoders: Regularizing Variational Autoencoder Latent Spaces
- Speech-to-speech Translation between Untranscribed Unknown Languages
- Neural-Symbolic Descriptive Action Model from Images: The Search for STRIPS
- High-Resolution Complex Scene Synthesis with Transformers
- DeepFracture: A Generative Approach for Predicting Brittle Fractures with Neural Discrete Representation Learning
- Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction
- GA-GAN: CT reconstruction from Biplanar DRRs using GAN with Guided Attention
- VQ-DRAW: A Sequential Discrete VAE
- Unsupervised Learning For Sequence-to-sequence Text-to-speech For Low-resource Languages
- Infomax Neural Joint Source-Channel Coding via Adversarial Bit Flip
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- Discond-VAE: Disentangling Continuous Factors from the Discrete
- Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders
- Compositional Fine-Grained Low-Shot Learning
- HAVANA: Hierarchical and Variation-Normalized Autoencoder for Person Re-identification
- Fast and Scalable Image Search For Histology
- Defending Against Image Corruptions Through Adversarial Augmentations
- CVC: Contrastive Learning for Non-parallel Voice Conversion
- UWSpeech: Speech to Speech Translation for Unwritten Languages
- Blur, Noise, and Compression Robust Generative Adversarial Networks
- crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder
- Hierarchical disentangled representation learning for singing voice conversion
- Decomposing Normal and Abnormal Features of Medical Images into Discrete Latent Codes for Content-Based Image Retrieval
- I'm Sorry for Your Loss: Spectrally-Based Audio Distances Are Bad at Pitch
- Nonparallel Voice Conversion with Augmented Classifier Star Generative Adversarial Networks
- Optimal Variance Control of the Score Function Gradient Estimator for Importance Weighted Bounds
- Exploring Fluent Query Reformulations with Text-to-Text Transformers and Reinforcement Learning
- not-so-BigGAN: Generating High-Fidelity Images on Small Compute with Wavelet-based Super-Resolution
- Semantics-Aware Human Motion Generation from Audio Instructions
- NASH: Toward End-to-End Neural Architecture for Generative Semantic Hashing
- The ELBO of Variational Autoencoders Converges to a Sum of Three Entropies
- Non-Intrusive Binaural Speech Intelligibility Prediction from Discrete Latent Representations
- One-Class SVM on siamese neural network latent space for Unsupervised Anomaly Detection on brain MRI White Matter Hyperintensities
- asya: Mindful verbal communication using deep learning
- Unsupervised Acoustic Unit Representation Learning for Voice Conversion using WaveNet Auto-encoders
- Learning Sampling in Financial Statement Audits using Vector Quantised Autoencoder Neural Networks
- MaGNET: Uniform Sampling from Deep Generative Network Manifolds Without Retraining
- Foveation for Segmentation of Ultra-High Resolution Images
- Likelihood Assignment for Out-of-Distribution Inputs in Deep Generative Models is Sensitive to Prior Distribution Choice
- Towards Interlingua Neural Machine Translation
- NWT: Towards natural audio-to-video generation with representation learning
- Robust Representation Learning via Perceptual Similarity Metrics
- Fast gradient-free activation maximization for neurons in spiking neural networks
- Set Distribution Networks: a Generative Model for Sets of Images
- Learning Product Codebooks using Vector Quantized Autoencoders for Image Retrieval
- Semi-Supervised Hierarchical Drug Embedding in Hyperbolic Space
- A Short Note on Analyzing Sequence Complexity in Trajectory Prediction Benchmarks
- Latent Space Optimal Transport for Generative Models
- Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning
- Unsupervised Multimodal Word Discovery based on Double Articulation Analysis with Co-occurrence cues
- Disentanglement of Latent Representations via Causal Interventions
- Unsupervised Source Separation via Bayesian Inference in the Latent Domain
- Direct Noisy Speech Modeling for Noisy-to-Noisy Voice Conversion
- Synthesising Activity Participations and Scheduling with Deep Generative Machine Learning
- Reconstructing hadronically decaying tau leptons with a jet foundation model
- Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices
- L-Verse: Bidirectional Generation Between Image and Text
- Discrete Acoustic Space for an Efficient Sampling in Neural Text-To-Speech
- The Image Local Autoregressive Transformer
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Trainable Class Prototypes for Few-Shot Learning
- Exploring Disentanglement with Multilingual and Monolingual VQ-VAE
- Autoencoding Under Normalization Constraints
- Abstract Reasoning via Logic-guided Generation
- Improving Lossless Compression Rates via Monte Carlo Bits-Back Coding
- Neural Bayes: A Generic Parameterization Method for Unsupervised Representation Learning
- Latent Vector Recovery of Audio GANs
- Incorporating Real-world Noisy Speech in Neural-network-based Speech Enhancement Systems
- A Framework for Generative and Contrastive Learning of Audio Representations
- Relaxed-Responsibility Hierarchical Discrete VAEs
- Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
- Reliable Categorical Variational Inference with Mixture of Discrete Normalizing Flows
- A Factorial Mixture Prior for Compositional Deep Generative Models
- Cauchy-Schwarz Regularized Autoencoder
- Neural Communication Systems with Bandwidth-limited Channel
- Learning to Generate 3D Shapes with Generative Cellular Automata
- Unsupervised Disentanglement of Linear-Encoded Facial Semantics
- HRINet: Alternative Supervision Network for High-resolution CT image Interpolation
- Incorporating Symbolic Sequential Modeling for Speech Enhancement
- No Representation without Transformation
- MaskAAE: Latent space optimization for Adversarial Auto-Encoders
- Many-to-Many Voice Conversion using Cycle-Consistent Variational Autoencoder with Multiple Decoders
- Discrete Few-Shot Learning for Pan Privacy
- Density Deconvolution with Normalizing Flows
- Generative networks as inverse problems with fractional wavelet scattering networks
- Hierarchical Timbre-Painting and Articulation Generation
- Generative Model without Prior Distribution Matching
- The Utility of Decorrelating Colour Spaces in Vector Quantised Variational Autoencoders
- Supervised Vector Quantized Variational Autoencoder for Learning Interpretable Global Representations
- Bridging the ELBO and MMD
- Feedback Recurrent AutoEncoder
- AE-OT-GAN: Training GANs from data specific latent distribution
- Variational Learning for Unsupervised Knowledge Grounded Dialogs
- Cross-modal Spectrum Transformation Network For Acoustic Scene classification
- Applying the Information Bottleneck Principle to Prosodic Representation Learning
- EdiBERT, a generative model for image editing
- Computational principles of intelligence: learning and reasoning with neural networks
- Analysis of ODE2VAE with Examples
- CNN with large memory layers
- Eccentric Regularization: Minimizing Hyperspherical Energy without explicit projection
- Layered Controllable Video Generation
- Discrete Variational Attention Models for Language Generation
- Causal Representation Learning for Context-Aware Face Transfer
- Measuring the Stability of Learned Features
- Introducing: DeepHead, Wide-band Electromagnetic Imaging Paradigm
- Unsupervised Skill-Discovery and Skill-Learning in Minecraft
- Synthetic weather radar using hybrid quantum-classical machine learning
- PocketVAE: A Two-step Model for Groove Generation and Control
- PixelTransformer: Sample Conditioned Signal Generation
- BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features
- Discrete representations in neural models of spoken language
- IB-DRR: Incremental Learning with Information-Back Discrete Representation Replay
- NP-DRAW: A Non-Parametric Structured Latent Variable Model for Image Generation
- Polyline Generative Navigable Space Segmentation for Autonomous Visual Navigation
- Depthwise Discrete Representation Learning
- Unsupervised Program Synthesis for Images By Sampling Without Replacement
- Inverse Graphics: Unsupervised Learning of 3D Shapes from Single Images
- VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space
- Learning Interpretable and Discrete Representations with Adversarial Training for Unsupervised Text Classification
- Complex-Valued Restricted Boltzmann Machine for Direct Speech Parameterization from Complex Spectra
- Learning source-aware representations of music in a discrete latent space
- Learned transform compression with optimized entropy encoding
- Voice Conversion Based Speaker Normalization for Acoustic Unit Discovery
- Noise Robust Generative Adversarial Networks
- Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 Challenge
- Supervised Encoding for Discrete Representation Learning
- Preliminary study on using vector quantization latent spaces for TTS/VC systems with consistent performance
- Discrete Auto-regressive Variational Attention Models for Text Modeling
- A New Framework for Machine Intelligence: Concepts and Prototype
- Evidence-Aware Inferential Text Generation with Vector Quantised Variational AutoEncoder
- Unsupervised Word Segmentation from Discrete Speech Units in Low-Resource Settings
- Learning the Imaging Landmarks: Unsupervised Key point Detection in Lung Ultrasound Videos
- Semi-supervised Grasp Detection by Representation Learning in a Vector Quantized Latent Space
- Lattice Representation Learning
- Information Bottleneck Approach to Spatial Attention Learning
- Quantization-Based Regularization for Autoencoders
- A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion
- GIFnets: Differentiable GIF Encoding Framework
- Word Representation for Rhythms
- Entropy optimized semi-supervised decomposed vector-quantized variational autoencoder model based on transfer learning for multiclass text classification and generation
- Comparative Snippet Generation
- Efficient And Scalable Neural Residual Waveform Coding With Collaborative Quantization
- Illiterate DALL-E Learns to Compose
- Synthesizing Photorealistic Images with Deep Generative Learning
- Zero-shot Voice Conversion via Self-supervised Prosody Representation Learning
- Momentum Contrastive Autoencoder: Using Contrastive Learning for Latent Space Distribution Matching in WAE
- Symbolic Music Loop Generation with VQ-VAE
- Deep Conditional Measure Quantization
- Variational latent discrete representation for time series modelling
- Exploration into Translation-Equivariant Image Quantization
- Unified Signal Compression Using a GAN with Iterative Latent Representation Optimization
- Learnable Triangulation for Deep Learning-based 3D Reconstruction of Objects of Arbitrary Topology from Single RGB Images
- Deep Sequence Learning for Video Anticipation: From Discrete and Deterministic to Continuous and Stochastic
- Truth-Conditional Captioning of Time Series Data
- High Mutual Information in Representation Learning with Symmetric Variational Inference
- Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm
- Learning Speaker Embedding from Text-to-Speech
- DiscoDVT: Generating Long Text with Discourse-Aware Discrete Variational Transformer
- A Comparison of Discrete Latent Variable Models for Speech Representation Learning
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- Recent Advancements in Self-Supervised Paradigms for Visual Feature Representation
- Learning from Multiple Time Series: A Deep Disentangled Approach to Diversified Time Series Forecasting
- Disentangling Generative Factors in Natural Language with Discrete Variational Autoencoders
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
- A Temporal Variational Model for Story Generation
- Testing for Typicality with Respect to an Ensemble of Learned Distributions
- Mutual Information Constraints for Monte-Carlo Objectives
- Physics Driven Domain Specific Transporter Framework with Attention Mechanism for Ultrasound Imaging
- Learning Physical Concepts in Cyber-Physical Systems: A Case Study
- StreamHover: Livestream Transcript Summarization and Annotation
- On the difficulty of a distributional semantics of spoken language
- A Generalised Linear Model Framework for -Variational Autoencoders based on Exponential Dispersion Families
- End-To-End Dilated Variational Autoencoder with Bottleneck Discriminative Loss for Sound Morphing -- A Preliminary Study
- Irregular Convolutional Auto-Encoder on Point Clouds
- Differentiable Segmentation of Sequences