Pay Less Attention with Lightweight and Dynamic Convolutions
arXiv:1901.10430
Abstract
Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that a very lightweight convolution can perform competitively to the best reported self-attention results. Next, we introduce dynamic convolutions which are simpler and more efficient than self-attention. We predict separate convolution kernels based solely on the current time-step in order to determine the importance of context elements. The number of operations required by this approach scales linearly in the input length, whereas self-attention is quadratic. Experiments on large-scale machine translation, language modeling and abstractive summarization show that dynamic convolutions improve over strong self-attention models. On the WMT'14 English-German test set dynamic convolutions achieve a new state of the art of 29.7 BLEU.
14 pages, ICLR oral
Cited by in corpus (137)
- BERTScore: Evaluating Text Generation with BERT
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- MLP-Mixer: An all-MLP Architecture for Vision
- Efficiently Modeling Long Sequences with Structured State Spaces
- Beyond English-Centric Multilingual Machine Translation
- Conformer: Convolution-augmented Transformer for Speech Recognition
- SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Reducing Transformer Depth on Demand with Structured Dropout
- HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
- Stand-Alone Self-Attention in Vision Models
- Incorporating BERT into Neural Machine Translation
- HiPPO: Recurrent Memory with Optimal Polynomial Projections
- Random Feature Attention
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- Training with Quantization Noise for Extreme Model Compression
- Lightweight Self-Attentive Sequential Recommendation
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- Very Deep Transformers for Neural Machine Translation
- Aligned Cross Entropy for Non-Autoregressive Machine Translation
- Non-Autoregressive Machine Translation with Disentangled Context Transformer
- Semi-Autoregressive Training Improves Mask-Predict Decoding
- DeLighT: Deep and Light-weight Transformer
- Augmenting Self-attention with Persistent Memory
- GNN-FiLM: Graph Neural Networks with Feature-wise Linear Modulation
- Data Diversification: A Simple Strategy For Neural Machine Translation
- MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning
- Container: Context Aggregation Network
- Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention
- Low-Memory Neural Network Training: A Technical Report
- Multi-modality Latent Interaction Network for Visual Question Answering
- Are Pre-trained Convolutions Better than Pre-trained Transformers?
- Exploring Self-attention for Image Recognition
- Pay Attention to MLPs
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
- Taming Sparsely Activated Transformer with Stochastic Experts
- Understanding the Difficulty of Training Transformers
- On Feature Normalization and Data Augmentation
- Tree-structured Attention with Hierarchical Accumulation
- Revisiting Low-Resource Neural Machine Translation: A Case Study
- TranSmart: A Practical Interactive Machine Translation System
- Learning Graph Structures with Transformer for Multivariate Time Series Anomaly Detection in IoT
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- Distilling Knowledge Learned in BERT for Text Generation
- Adaptively Sparse Transformers
- Joint Source-Target Self Attention with Locality Constraints
- GNN-LM: Language Modeling based on Global Contexts via GNN
- Visual Parser: Representing Part-whole Hierarchies with Transformers
- Strategies for Structuring Story Generation
- PowerNorm: Rethinking Batch Normalization in Transformers
- Self-Knowledge Distillation with Progressive Refinement of Targets
- Time-aware Large Kernel Convolutions
- Multi-branch Attentive Transformer
- On Exposure Bias, Hallucination and Domain Shift in Neural Machine Translation
- Fast Convergence of DETR with Spatially Modulated Co-Attention
- Cross-model Back-translated Distillation for Unsupervised Machine Translation
- Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling
- Normalized Attention Without Probability Cage
- Hard-Coded Gaussian Attention for Neural Machine Translation
- Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
- POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-training
- Neural Machine Translation: Challenges, Progress and Future
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Language Models are Good Translators
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- Positional Normalization
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Mask Attention Networks: Rethinking and Strengthen Transformer
- Recurrent Graph Syntax Encoder for Neural Machine Translation
- Self-Attention Enhanced Selective Gate with Entity-Aware Embedding for Distantly Supervised Relation Extraction
- On Infinite-Width Hypernetworks
- DLA-Net: Learning Dual Local Attention Features for Semantic Segmentation of Large-Scale Building Facade Point Clouds
- Abstractive Text Summarization based on Language Model Conditioning and Locality Modeling
- Data Rejuvenation: Exploiting Inactive Training Examples for Neural Machine Translation
- Dual Contrastive Loss and Attention for GANs
- Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention
- Not All Memories are Created Equal: Learning to Forget by Expiring
- Learning Language Specific Sub-network for Multilingual Machine Translation
- FastFusionNet: New State-of-the-Art for DAWNBench SQuAD
- Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size
- Fast Interleaved Bidirectional Sequence Generation
- Multi-Unit Transformers for Neural Machine Translation
- Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning
- MUFASA: Multimodal Fusion Architecture Search for Electronic Health Records
- Axial Residual Networks for CycleGAN-based Voice Conversion
- Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine Translation
- Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention
- IOT: Instance-wise Layer Reordering for Transformer Structures
- Towards Enhancing Database Education: Natural Language Generation Meets Query Execution Plans
- A Unified Neural Coherence Model
- A Lightweight dynamic filter for keyword spotting
- On the Importance of Local Information in Transformer Based Models
- Neural Machine Translation: A Review and Survey
- Shallow-to-Deep Training for Neural Machine Translation
- An Attention Module for Convolutional Neural Networks
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation
- GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
- Learning Light-Weight Translation Models from Deep Transformer
- Improving Neural Machine Translation by Bidirectional Training
- Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers
- Language Models not just for Pre-training: Fast Online Neural Noisy Channel Modeling
- Attention-based ASR with Lightweight and Dynamic Convolutions
- Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models
- Deeper or Wider Networks of Point Clouds with Self-attention?
- Probabilistic Attention for Interactive Segmentation
- Training Flexible Depth Model by Multi-Task Learning for Neural Machine Translation
- Multi-Task Sequence Prediction For Tunisian Arabizi Multi-Level Annotation
- Iterative Batch Back-Translation for Neural Machine Translation: A Conceptual Model
- UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost
- Hierarchical Attention Transformer Architecture For Syntactic Spell Correction
- Do Transformers Need Deep Long-Range Memory
- Context-Gated Convolution
- Dynamic Convolution for 3D Point Cloud Instance Segmentation
- Reciprocal Supervised Learning Improves Neural Machine Translation
- RankNAS: Efficient Neural Architecture Search by Pairwise Ranking
- The NiuTrans System for WNGT 2020 Efficiency Task
- Low-Resource Dialogue Summarization with Domain-Agnostic Multi-Source Pretraining
- Topic-Aware Contrastive Learning for Abstractive Dialogue Summarization
- UdS Submission for the WMT 19 Automatic Post-Editing Task
- The University of Sydney's Machine Translation System for WMT19
- Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text Generation
- Normalization of Input-output Shared Embeddings in Text Generation Models
- Resurrecting Submodularity for Neural Text Generation
- Seven Myths in Machine Learning Research
- The Volctrans Neural Speech Translation System for IWSLT 2021
- Langsmith: An Interactive Academic Text Revision System
- Dialogue Summarization with Supporting Utterance Flow Modeling and Fact Regularization
- On the Sparsity of Neural Machine Translation Models
- Two-Headed Monster And Crossed Co-Attention Networks
- Dynamic Parameterized Network for CTR Prediction
- FastTrees: Parallel Latent Tree-Induction for Faster Sequence Encoding
- Layer-Wise Multi-View Learning for Neural Machine Translation
- Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
- Higher-order Network for Action Recognition
- Empirical Analysis of Korean Public AI Hub Parallel Corpora and in-depth Analysis using LIWC
- Revisit Systematic Generalization via Meaningful Learning