Effective Approaches to Attention-based Neural Machine Translation
arXiv:1508.04025
Abstract
An attentional mechanism has lately been used to improve neural machine translation (NMT) by selectively focusing on parts of the source sentence during translation. However, there has been little work exploring useful architectures for attention-based NMT. This paper examines two simple and effective classes of attentional mechanism: a global approach which always attends to all source words and a local one that only looks at a subset of source words at a time. We demonstrate the effectiveness of both approaches over the WMT translation tasks between English and German in both directions. With local attention, we achieve a significant gain of 5.0 BLEU points over non-attentional systems which already incorporate known techniques such as dropout. Our ensemble model using different attention architectures has established a new state-of-the-art result in the WMT'15 English to German translation task with 25.9 BLEU points, an improvement of 1.0 BLEU points over the existing best system backed by NMT and an n-gram reranker.
11 pages, 7 figures, EMNLP 2015 camera-ready version, more training details
References in corpus (3)
Cited by in corpus (261)
- Convolutional Sequence to Sequence Learning
- DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
- Attention in Natural Language Processing
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- Understanding Neural Networks through Representation Erasure
- The Natural Language Decathlon: Multitask Learning as Question Answering
- Quasi-Recurrent Neural Networks
- Escaping the Big Data Paradigm with Compact Transformers
- Selective Encoding for Abstractive Sentence Summarization
- A Survey of Knowledge-Enhanced Text Generation
- Online and Linear-Time Attention by Enforcing Monotonic Alignments
- Transformers for Modeling Physical Systems
- Modeling Coverage for Neural Machine Translation
- A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task
- Language to Logical Form with Neural Attention
- DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
- Sequence-Level Knowledge Distillation
- Hierarchical Generation of Molecular Graphs using Structural Motifs
- Generating News Headlines with Recurrent Neural Networks
- Learning Longer-term Dependencies in RNNs with Auxiliary Losses
- Tree-to-tree Neural Networks for Program Translation
- OpenNMT: Neural Machine Translation Toolkit
- Aspect Level Sentiment Classification with Deep Memory Network
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language Translation
- Deeply learning molecular structure-property relationships using attention- and gate-augmented graph convolutional network
- Sparse Sinkhorn Attention
- Learning to Represent Edits
- KERMIT: Generative Insertion-Based Modeling for Sequences
- Neural Machine Translation with Reconstruction
- Multilingual Extractive Reading Comprehension by Runtime Machine Translation
- Towards better decoding and language model integration in sequence to sequence models
- IncSQL: Training Incremental Text-to-SQL Parsers with Non-Deterministic Oracles
- DP-GAN: Diversity-Promoting Generative Adversarial Network for Generating Informative and Diversified Text
- Analyzing Uncertainty in Neural Machine Translation
- Challenges in Data-to-Document Generation
- A Spatio-Temporal Spot-Forecasting Framework for Urban Traffic Prediction
- SGM: Sequence Generation Model for Multi-label Classification
- Multilingual Hierarchical Attention Networks for Document Classification
- S-Net: From Answer Extraction to Answer Generation for Machine Reading Comprehension
- Bridging Neural Machine Translation and Bilingual Dictionaries
- A neural network walks into a lab: towards using deep nets as models for human behavior
- Domain specialization: a post-training domain adaptation for Neural Machine Translation
- Controlling Output Length in Neural Encoder-Decoders
- Minimum Risk Training for Neural Machine Translation
- Tree-to-Sequence Attentional Neural Machine Translation
- Towards Binary-Valued Gates for Robust LSTM Training
- Speech Emotion Recognition via Contrastive Loss under Siamese Networks
- Vocabulary Selection Strategies for Neural Machine Translation
- Dank Learning: Generating Memes Using Deep Neural Networks
- A Transformer-based Approach for Source Code Summarization
- Point2Sequence: Learning the Shape Representation of 3D Point Clouds with an Attention-based Sequence to Sequence Network
- Coverage Embedding Models for Neural Machine Translation
- Bridging the Gap between Spatial and Spectral Domains: A Survey on Graph Neural Networks
- The computerization of archaeology: survey on AI techniques
- Variational Neural Machine Translation
- An attentive neural architecture for joint segmentation and parsing and its application to real estate ads
- Aspect Term Extraction with History Attention and Selective Transformation
- Exploiting Cross-Sentence Context for Neural Machine Translation
- Joint Training for Neural Machine Translation Models with Monolingual Data
- Neural Machine Translation with Pivot Languages
- Text normalization using memory augmented neural networks
- Supervised Attentions for Neural Machine Translation
- Self-Supervised Gait Encoding with Locality-Aware Attention for Person Re-Identification
- AutoLoss: Learning Discrete Schedules for Alternate Optimization
- Context-Aware Self-Attention Networks
- An Attentional Neural Conversation Model with Improved Specificity
- The Impact of Preprocessing on Arabic-English Statistical and Neural Machine Translation
- A GRU-Gated Attention Model for Neural Machine Translation
- Neural Machine Translation with Supervised Attention
- A Deep Memory-based Architecture for Sequence-to-Sequence Learning
- Sequence-to-Sequence Data Augmentation for Dialogue Language Understanding
- Seq2Seq AI Chatbot with Attention Mechanism
- High Quality Prediction of Protein Q8 Secondary Structure by Diverse Neural Network Architectures
- Data Recombination for Neural Semantic Parsing
- Exploiting Deep Representations for Neural Machine Translation
- Accelerating Neural Transformer via an Average Attention Network
- Context in Neural Machine Translation: A Review of Models and Evaluations
- Improving Graph Neural Network Representations of Logical Formulae with Subgraph Pooling
- Distance Metric Learning for Aspect Phrase Grouping
- Global Encoding for Abstractive Summarization
- Learning to Mine Aligned Code and Natural Language Pairs from Stack Overflow
- Iterative Refinement for Machine Translation
- Learn to Code-Switch: Data Augmentation using Copy Mechanism on Language Modeling
- A Bi-LSTM-RNN Model for Relation Classification Using Low-Cost Sequence Features
- Sequential Context Encoding for Duplicate Removal
- Learning to Parse and Translate Improves Neural Machine Translation
- A Multi-task Learning Approach for Improving Product Title Compression with User Search Log Data
- Wat zei je? Detecting Out-of-Distribution Translations with Variational Transformers
- Semi-Supervised Sequence Modeling with Cross-View Training
- A Semantic Relevance Based Neural Network for Text Summarization and Text Simplification
- Self-Attentive Residual Decoder for Neural Machine Translation
- Neural Discourse Relation Recognition with Semantic Memory
- Modelling Sentence Pairs with Tree-structured Attentive Encoder
- Domain Control for Neural Machine Translation
- Real-world Ride-hailing Vehicle Repositioning using Deep Reinforcement Learning
- Visual Question Answering with Memory-Augmented Networks
- Deconvolution-Based Global Decoding for Neural Machine Translation
- Deep Neural Machine Translation with Linear Associative Unit
- Towards Coherent and Engaging Spoken Dialog Response Generation Using Automatic Conversation Evaluators
- A Question Answering Approach to Emotion Cause Extraction
- MAT: A Multimodal Attentive Translator for Image Captioning
- Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images
- Towards End-to-end Automatic Code-Switching Speech Recognition
- Hybrid Self-Attention Network for Machine Translation
- Source-side Prediction for Neural Headline Generation
- Learning to Remember Translation History with a Continuous Cache
- Variational Recurrent Neural Machine Translation
- Deep Neural Machine Translation with Weakly-Recurrent Units
- A Full End-to-End Semantic Role Labeler, Syntax-agnostic Over Syntax-aware?
- Deep Graph Attention Model
- OBJ2TEXT: Generating Visually Descriptive Language from Object Layouts
- Mining Significant Microblogs for Misinformation Identification: An Attention-based Approach
- Sentence Simplification with Deep Reinforcement Learning
- Improved Neural Text Attribute Transfer with Non-parallel Data
- Don't Classify, Translate: Multi-Level E-Commerce Product Categorization Via Machine Translation
- Phrase-Based Attentions
- CoCoSum: Contextual Code Summarization with Multi-Relational Graph Neural Network
- Navigating with Graph Representations for Fast and Scalable Decoding of Neural Language Models
- Positional Encoding to Control Output Sequence Length
- Learning Edge Properties in Graphs from Path Aggregations
- Autoencoder as Assistant Supervisor: Improving Text Representation for Chinese Social Media Text Summarization
- DeepHTTP: Semantics-Structure Model with Attention for Anomalous HTTP Traffic Detection and Pattern Mining
- Operations Guided Neural Networks for High Fidelity Data-To-Text Generation
- Neural Machine Translation with Key-Value Memory-Augmented Attention
- Online Segment to Segment Neural Transduction
- Double Path Networks for Sequence to Sequence Learning
- Still not there? Comparing Traditional Sequence-to-Sequence Models to Encoder-Decoder Neural Networks on Monotone String Translation Tasks
- Machine Translation at Booking.com: Journey and Lessons Learned
- Lexicons and Minimum Risk Training for Neural Machine Translation: NAIST-CMU at WAT2016
- Unsupervised Neural Hidden Markov Models
- Toward a full-scale neural machine translation in production: the Booking.com use case
- Language Graph Distillation for Low-Resource Machine Translation
- Hard but Robust, Easy but Sensitive: How Encoder and Decoder Perform in Neural Machine Translation
- Deriving Machine Attention from Human Rationales
- Gated Recurrent Context: Softmax-free Attention for Online Encoder-Decoder Speech Recognition
- seq2graph: Discovering Dynamic Dependencies from Multivariate Time Series with Multi-level Attention
- Effective writing style imitation via combinatorial paraphrasing
- Guiding attention in Sequence-to-sequence models for Dialogue Act prediction
- Learning Autocomplete Systems as a Communication Game
- Sharing Attention Weights for Fast Transformer
- Structured-based Curriculum Learning for End-to-end English-Japanese Speech Translation
- A study of latent monotonic attention variants
- Why not be Versatile? Applications of the SGNMT Decoder for Machine Translation
- Structurally Sparsified Backward Propagation for Faster Long Short-Term Memory Training
- Same Representation, Different Attentions: Shareable Sentence Representation Learning from Multiple Tasks
- BPE and CharCNNs for Translation of Morphology: A Cross-Lingual Comparison and Analysis
- Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models
- Differentiable lower bound for expected BLEU score
- Improving Textual Network Embedding with Global Attention via Optimal Transport
- SuperChat: Dialogue Generation by Transfer Learning from Vision to Language using Two-dimensional Word Embedding and Pretrained ImageNet CNN Models
- Deep Learning for UL/DL Channel Calibration in Generic Massive MIMO Systems
- An Auto-Encoder Matching Model for Learning Utterance-Level Semantic Dependency in Dialogue Generation
- Sequential Copying Networks
- Automatically Generating Commit Messages from Diffs using Neural Machine Translation
- ACE-NODE: Attentive Co-Evolving Neural Ordinary Differential Equations
- Pattern Generation Strategies for Improving Recognition of Handwritten Mathematical Expressions
- Progress and Tradeoffs in Neural Language Models
- Deep Neural Network for Semantic-based Text Recognition in Images
- Adversarial Domain Adaptation for Duplicate Question Detection
- A Stable and Effective Learning Strategy for Trainable Greedy Decoding
- Going Wider: Recurrent Neural Network With Parallel Cells
- Learning from Fact-checkers: Analysis and Generation of Fact-checking Language
- Complexity-Weighted Loss and Diverse Reranking for Sentence Simplification
- Improving Semantic Relevance for Sequence-to-Sequence Learning of Chinese Social Media Text Summarization
- Teaching Machines to Code: Neural Markup Generation with Visual Attention
- Spatio-Temporal Point Processes with Attention for Traffic Congestion Event Modeling
- Distilling Knowledge for Search-based Structured Prediction
- "Found in Translation": Predicting Outcomes of Complex Organic Chemistry Reactions using Neural Sequence-to-Sequence Models
- Learning to Discriminate Noises for Incorporating External Information in Neural Machine Translation
- Attention Models for Point Clouds in Deep Learning: A Survey
- TransOMCS: From Linguistic Graphs to Commonsense Knowledge
- Upcycle Your OCR: Reusing OCRs for Post-OCR Text Correction in Romanised Sanskrit
- Learning to Create Better Ads: Generation and Ranking Approaches for Ad Creative Refinement
- Augmenting Neural Networks with First-order Logic
- Fusing Recency into Neural Machine Translation with an Inter-Sentence Gate Model
- Restoring ancient text using deep learning: a case study on Greek epigraphy
- Aiming to Know You Better Perhaps Makes Me a More Engaging Dialogue Partner
- Memory-Augmented Neural Networks for Machine Translation
- Title-Guided Encoding for Keyphrase Generation
- Neural Machine Translation via Binary Code Prediction
- Semantic-Unit-Based Dilated Convolution for Multi-Label Text Classification
- Attentive Convolution: Equipping CNNs with RNN-style Attention Mechanisms
- Direct Output Connection for a High-Rank Language Model
- Why and How to Pay Different Attention to Phrase Alignments of Different Intensities
- Recursive Neural Network Based Preordering for English-to-Japanese Machine Translation
- Document Graph for Neural Machine Translation
- Multi-Horizon Forecasting for Limit Order Books: Novel Deep Learning Approaches and Hardware Acceleration using Intelligent Processing Units
- Influence-aware Memory Architectures for Deep Reinforcement Learning
- Syntax-based Attention Model for Natural Language Inference
- Multi-Stream End-to-End Speech Recognition
- Target Guided Emotion Aware Chat Machine
- Learning Comment Generation by Leveraging User-Generated Data
- Inducing Grammars with and for Neural Machine Translation
- A Correlational Encoder Decoder Architecture for Pivot Based Sequence Generation
- Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
- End-to-End Argument Mining for Discussion Threads Based on Parallel Constrained Pointer Architecture
- GILE: A Generalized Input-Label Embedding for Text Classification
- Exact-K Recommendation via Maximal Clique Optimization
- Syntax-Enhanced Neural Machine Translation with Syntax-Aware Word Representations
- A Deep Learning Approach to Automate High-Resolution Blood Vessel Reconstruction on Computerized Tomography Images With or Without the Use of Contrast Agent
- Generating Dialogue Responses from a Semantic Latent Space
- Neural Machine Translation Model with a Large Vocabulary Selected by Branching Entropy
- Modeling Past and Future for Neural Machine Translation
- Video Question Generation via Cross-Modal Self-Attention Networks Learning
- Choose Your Programming Copilot: A Comparison of the Program Synthesis Performance of GitHub Copilot and Genetic Programming
- Temporal Attention-Gated Model for Robust Sequence Classification
- Towards Enhancing Database Education: Natural Language Generation Meets Query Execution Plans
- Transformers are Deep Infinite-Dimensional Non-Mercer Binary Kernel Machines
- Generating Semantically Valid Adversarial Questions for TableQA
- Back-Translation Sampling by Targeting Difficult Words in Neural Machine Translation
- Iterative Batch Back-Translation for Neural Machine Translation: A Conceptual Model
- Exploring the Use of Attention within an Neural Machine Translation Decoder States to Translate Idioms
- Experiential Robot Learning with Accelerated Neuroevolution
- Towards User Friendly Medication Mapping Using Entity-Boosted Two-Tower Neural Network
- Image Captioning as Neural Machine Translation Task in SOCKEYE
- Multi-scale Alignment and Contextual History for Attention Mechanism in Sequence-to-sequence Model
- Neural Morphological Tagging for Estonian
- When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
- Context-aware Cascade Attention-based RNN for Video Emotion Recognition
- Pointing to Subwords for Generating Function Names in Source Code
- KNPTC: Knowledge and Neural Machine Translation Powered Chinese Pinyin Typo Correction
- Real-time low-resource phoneme recognition on edge devices
- Exploring Neural Methods for Parsing Discourse Representation Structures
- PLANS: Robust Program Learning from Neurally Inferred Specifications
- Language-Independent Representor for Neural Machine Translation
- Generating Diverse Translation by Manipulating Multi-Head Attention
- Incorporating Textual Evidence in Visual Storytelling
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
- A Comprehensive Study on Temporal Modeling for Online Action Detection
- Evaluating Sequence-to-Sequence Learning Models for If-Then Program Synthesis
- Implicit Argument Prediction as Reading Comprehension
- "Wait, I'm Still Talking!" Predicting the Dialogue Interaction Behavior Using Imagine-Then-Arbitrate Model
- CAESAR: Context Awareness Enabled Summary-Attentive Reader
- Learning When to Concentrate or Divert Attention: Self-Adaptive Attention Temperature for Neural Machine Translation
- Controlling Neural Machine Translation Formality with Synthetic Supervision
- Hierarchical RNN with Static Sentence-Level Attention for Text-Based Speaker Change Detection
- Beyond Weight Tying: Learning Joint Input-Output Embeddings for Neural Machine Translation
- Improving Scientific Article Visibility by Neural Title Simplification
- Neural Sequence Model Training via -divergence Minimization
- Recovery command generation towards automatic recovery in ICT systems by Seq2Seq learning
- Automatic Documentation of ICD Codes with Far-Field Speech Recognition
- GRET: Global Representation Enhanced Transformer
- Sent2Matrix: Folding Character Sequences in Serpentine Manifolds for Two-Dimensional Sentence
- Towards Fluent Translations from Disfluent Speech
- Future-Prediction-Based Model for Neural Machine Translation
- VideoMCC: a New Benchmark for Video Comprehension
- Modeling Homophone Noise for Robust Neural Machine Translation
- Graph-based Filtering of Out-of-Vocabulary Words for Encoder-Decoder Models
- Can DNNs Learn to Lipread Full Sentences?
- Data Ordering Patterns for Neural Machine Translation: An Empirical Study
- Predicting In-game Actions from Interviews of NBA Players
- Phonetic-and-Semantic Embedding of Spoken Words with Applications in Spoken Content Retrieval
- A Sequence-to-Sequence Model for Semantic Role Labeling
- Pre-train, Interact, Fine-tune: A Novel Interaction Representation for Text Classification
- EAT: a simple and versatile semantic representation format for multi-purpose NLP
- Improving Target-side Lexical Transfer in Multilingual Neural Machine Translation
- Automatically Generating Codes from Graphical Screenshots Based on Deep Autocoder
- Differentiable Window for Dynamic Local Attention
- Gated Attentive-Autoencoder for Content-Aware Recommendation
- Quantum Statistics-Inspired Neural Attention