Language Modeling with Gated Convolutional Networks
arXiv:1612.08083
Abstract
The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a finite context approach through stacked convolutions, which can be more efficient since they allow parallelization over sequential tokens. We propose a novel simplified gating mechanism that outperforms Oord et al (2016) and investigate the impact of key architectural decisions. The proposed approach achieves state-of-the-art on the WikiText-103 benchmark, even though it features long-term dependencies, as well as competitive results on the Google Billion Words benchmark. Our model reduces the latency to score a sentence by an order of magnitude compared to a recurrent baseline. To our knowledge, this is the first time a non-recurrent approach is competitive with strong recurrent models on these large scale language tasks.
References in corpus (3)
Cited by in corpus (128)
- Convolutional Sequence to Sequence Learning
- Recent Trends in Deep Learning Based Natural Language Processing
- Comparative Study of CNN and RNN for Natural Language Processing
- Searching for Activation Functions
- Attention-based Deep Multiple Instance Learning
- Whole Slide Images based Cancer Survival Prediction using Attention Guided Deep Multiple Instance Learning Networks
- A Literature Survey of Recent Advances in Chatbots
- Parallel Spatio-Temporal Attention-Based TCN for Multivariate Time Series Prediction
- Jointly Multiple Events Extraction via Attention-based Graph Information Aggregation
- Hybrid Quantum-Classical Convolutional Neural Networks
- Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
- Learnable pooling with Context Gating for video classification
- Heterogeneous Graph Contrastive Learning for Recommendation
- Language Modeling with Deep Transformers
- Multi-level Convolutional Autoencoder Networks for Parametric Prediction of Spatio-temporal Dynamics
- Multi-View Graph Convolutional Network for Multimedia Recommendation
- Modeling Human Motion with Quaternion-based Neural Networks
- GRNN: Generative Regression Neural Network -- A Data Leakage Attack for Federated Learning
- Deceiving End-to-End Deep Learning Malware Detectors using Adversarial Examples
- Adversarial Generation of Natural Language
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Deep Enhanced Representation for Implicit Discourse Relation Recognition
- Autoregressive Convolutional Neural Networks for Asynchronous Time Series
- An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
- Improving Variational Auto-Encoders using Householder Flow
- Representation Learning for Natural Language Processing
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network Training and Inference
- Nonparallel Emotional Speech Conversion
- Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling
- Lip-reading with Densely Connected Temporal Convolutional Networks
- A Survey of Deep Learning Methods for Relation Extraction
- CNN+CNN: Convolutional Decoders for Image Captioning
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- Deconvolutional Paragraph Representation Learning
- Exploring Human Mobility for Multi-Pattern Passenger Prediction: A Graph Learning Framework
- Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context
- Unsupervised Time Series Outlier Detection with Diversity-Driven Convolutional Ensembles -- Extended Version
- T-former: An Efficient Transformer for Image Inpainting
- VAE with a VampPrior
- Reading Scene Text with Attention Convolutional Sequence Modeling
- Forecast Network-Wide Traffic States for Multiple Steps Ahead: A Deep Learning Approach Considering Dynamic Non-Local Spatial Correlation and Non-Stationary Temporal Dependency
- FairSR: Fairness-aware Sequential Recommendation through Multi-Task Learning with Preference Graph Embeddings
- Compressive Transformers for Long-Range Sequence Modelling
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- 6GCVAE: Gated Convolutional Variational Autoencoder for IPv6 Target Generation
- CaloMan: Fast generation of calorimeter showers with density estimation on learned manifolds
- Technical Q&A Site Answer Recommendation via Question Boosting
- Stochastic Restoration of Heavily Compressed Musical Audio using Generative Adversarial Networks
- Deformable PV-RCNN: Improving 3D Object Detection with Learned Deformations
- Non-autoregressive Transformer-based End-to-end ASR using BERT
- Who Needs Words? Lexicon-Free Speech Recognition
- Encoding Sentences with Graph Convolutional Networks for Semantic Role Labeling
- Relational recurrent neural networks
- A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization
- Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling
- TSNAT: Two-Step Non-Autoregressvie Transformer Models for Speech Recognition
- Fake News Detection as Natural Language Inference
- Probabilistic Deep Learning to Quantify Uncertainty in Air Quality Forecasting
- Metapopulation Graph Neural Networks: Deep Metapopulation Epidemic Modeling with Human Mobility
- Fast Parametric Learning with Activation Memorization
- Highrisk Prediction from Electronic Medical Records via Deep Attention Networks
- Few-shot Learning with Meta Metric Learners
- GLU Variants Improve Transformer
- A Deep Generative Model of Speech Complex Spectrograms
- Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised Clustering
- Spatiotemporal information conversion machine for time-series prediction
- Learning Convolutional Text Representations for Visual Question Answering
- Have convolutions already made recurrence obsolete for unconstrained handwritten text recognition ?
- Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition
- Automating Pharmacovigilance Evidence Generation: Using Large Language Models to Produce Context-Aware SQL
- Multimodal Urban Sound Tagging with Spatiotemporal Context
- Large-scale weakly supervised audio classification using gated convolutional neural network
- A Scalable Framework for Automatic Playlist Continuation on Music Streaming Services
- A Neural Language Model for Dynamically Representing the Meanings of Unknown Words and Entities in a Discourse
- Online Phase Reconstruction via DNN-based Phase Differences Estimation
- An Attention-Based Word-Level Interaction Model: Relation Detection for Knowledge Base Question Answering
- Visual Text Correction
- Using Graph Neural Networks for Mass Spectrometry Prediction
- Detecting Frames in News Headlines and Lead Images in U.S. Gun Violence Coverage
- Generative Convolution Layer for Image Generation
- In Conclusion Not Repetition: Comprehensive Abstractive Summarization With Diversified Attention Based On Determinantal Point Processes
- A Gap-Based Framework for Chinese Word Segmentation via Very Deep Convolutional Networks
- CoLLIE: Continual Learning of Language Grounding from Language-Image Embeddings
- NAS-VAD: Neural Architecture Search for Voice Activity Detection
- EnGN: A High-Throughput and Energy-Efficient Accelerator for Large Graph Neural Networks
- Lightweight Adaptive Mixture of Neural and N-gram Language Models
- SemSegDepth: A Combined Model for Semantic Segmentation and Depth Completion
- Heterogeneous Data Fusion Considering Spatial Correlations using Graph Convolutional Networks and its Application in Air Quality Prediction
- Weighted Sigmoid Gate Unit for an Activation Function of Deep Neural Network
- Tag-less Back-Translation
- Variational Autoencoder with Implicit Optimal Priors
- StarGAN-ZSVC: Towards Zero-Shot Voice Conversion in Low-Resource Contexts
- Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems
- Contextual Lensing of Universal Sentence Representations
- A Study of the Plausibility of Attention between RNN Encoders in Natural Language Inference
- An Investigation of the Effectiveness of Phase for Audio Classification
- Convergence of a Relaxed Variable Splitting Method for Learning Sparse Neural Networks via , and transformed- Penalties
- Tensorized Self-Attention: Efficiently Modeling Pairwise and Global Dependencies Together
- End-to-End Text Classification via Image-based Embedding using Character-level Networks
- Time Distributed Deep Learning Models for Purely Exogenous Forecasting: Application to Water Table Depth Prediction using Weather Image Time Series
- Curb Your Carbon Emissions: Benchmarking Carbon Emissions in Machine Translation
- AttS2S-VC: Sequence-to-Sequence Voice Conversion with Attention and Context Preservation Mechanisms
- Neural Random Projections for Language Modelling
- Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention
- Deep Griffin-Lim Iteration
- CNNs, LSTMs, and Attention Networks for Pathology Detection in Medical Data
- A Survey on Awesome Korean NLP Datasets
- On the Privacy Risks of Deploying Recurrent Neural Networks in Machine Learning Models
- Title-Guided Encoding for Keyphrase Generation
- Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
- Rapid Parameter Estimation for Merging Massive Black Hole Binaries Using Continuous Normalizing Flows
- Uncertainty Intervals for Graph-based Spatio-Temporal Traffic Prediction
- Character n-gram Embeddings to Improve RNN Language Models
- Relevance-Promoting Language Model for Short-Text Conversation
- SAM-GCNN: A Gated Convolutional Neural Network with Segment-Level Attention Mechanism for Home Activity Monitoring
- Machine Translation between Vietnamese and English: an Empirical Study
- Speech Paralinguistic Approach for Detecting Dementia Using Gated Convolutional Neural Network
- DBNet: A Dual-branch Network Architecture Processing on Spectrum and Waveform for Single-channel Speech Enhancement
- Language Modeling with Highway LSTM
- Identifying Query-Relevant Neurons in Large Language Models for Long-Form Texts
- Cosine-similarity penalty to discriminate sound classes in weakly-supervised sound event detection
- A More Efficient Chinese Named Entity Recognition base on BERT and Syntactic Analysis
- Large Language Models -- the Future of Fundamental Physics?
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting
- Sound event detection using weakly-labeled semi-supervised data with GCRNNS, VAT and Self-Adaptive Label Refinement
- SAM: Semantic Attribute Modulation for Language Modeling and Style Variation
- Language modeling with Neural trans-dimensional random fields
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions