Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
arXiv:1609.08144
Abstract
Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive both in training and in translation inference. Also, most NMT systems have difficulty with rare words. These issues have hindered NMT's use in practical deployments and services, where both accuracy and speed are essential. In this work, we present GNMT, Google's Neural Machine Translation system, which attempts to address many of these issues. Our model consists of a deep LSTM network with 8 encoder and 8 decoder layers using attention and residual connections. To improve parallelism and therefore decrease training time, our attention mechanism connects the bottom layer of the decoder to the top layer of the encoder. To accelerate the final translation speed, we employ low-precision arithmetic during inference computations. To improve handling of rare words, we divide words into a limited set of common sub-word units ("wordpieces") for both input and output. This method provides a good balance between the flexibility of "character"-delimited models and the efficiency of "word"-delimited models, naturally handles translation of rare words, and ultimately improves the overall accuracy of the system. Our beam search technique employs a length-normalization procedure and uses a coverage penalty, which encourages generation of an output sentence that is most likely to cover all the words in the source sentence. On the WMT'14 English-to-French and English-to-German benchmarks, GNMT achieves competitive results to state-of-the-art. Using a human side-by-side evaluation on a set of isolated simple sentences, it reduces translation errors by an average of 60% compared to Google's phrase-based production system.
Cited by in corpus (654)
- Photonics for artificial intelligence and neuromorphic computing
- fastai: A Layered API for Deep Learning
- ERNIE: Enhanced Representation through Knowledge Integration
- StarCraft II: A New Challenge for Reinforcement Learning
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic Classification
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Quasi-Recurrent Neural Networks
- Neural Rating Regression with Abstractive Tips Generation for Recommendation
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Depthwise Separable Convolutions for Neural Machine Translation
- Learning to Remember Rare Events
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- An Actor-Critic Algorithm for Sequence Prediction
- Device Placement Optimization with Reinforcement Learning
- Stand-Alone Self-Attention in Vision Models
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
- A Deep Reinforcement Learning Chatbot
- Incorporating BERT into Neural Machine Translation
- MLPerf Training Benchmark
- Recurrent Neural Networks (RNNs): A gentle Introduction and Overview
- ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
- Is Neural Machine Translation Ready for Deployment? A Case Study on 30 Translation Directions
- Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation
- SignBERT+: Hand-model-aware Self-supervised Pre-training for Sign Language Understanding
- Weighted Transformer Network for Machine Translation
- Learning the PE Header, Malware Detection with Minimal Domain Knowledge
- Linking the Neural Machine Translation and the Prediction of Organic Chemistry Reactions
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- On the Explainability of Natural Language Processing Deep Models
- On the Dimensionality of Word Embedding
- Measurement of Anomalous Diffusion Using Recurrent Neural Networks
- Compressing Recurrent Neural Network with Tensor Train
- Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault Attacks
- A Survey of Deep Learning Techniques for Neural Machine Translation
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- Atomic Convolutional Networks for Predicting Protein-Ligand Binding Affinity
- Fluency-Guided Cross-Lingual Image Captioning
- Learning Deep Transformer Models for Machine Translation
- Block-Sparse Recurrent Neural Networks
- DEPTWEET: A Typology for Social Media Texts to Detect Depression Severities
- Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges
- DCN+: Mixed Objective and Deep Residual Coattention for Question Answering
- Pre-Training BERT on Arabic Tweets: Practical Considerations
- Optimization and Abstraction: A Synergistic Approach for Analyzing Neural Network Robustness
- Story Generation from Sequence of Independent Short Descriptions
- Regularizing Deep Neural Networks by Noise: Its Interpretation and Optimization
- Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs
- Generating Wikipedia by Summarizing Long Sequences
- SYSTRAN's Pure Neural Machine Translation Systems
- Reversible Architectures for Arbitrarily Deep Residual Neural Networks
- Exploring Neural Transducers for End-to-End Speech Recognition
- On the Effectiveness of Mitigating Data Poisoning Attacks with Gradient Shaping
- Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
- Massive Exploration of Neural Machine Translation Architectures
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- A Study of BFLOAT16 for Deep Learning Training
- Semantic Tagging with Deep Residual Networks
- Fast Structured Decoding for Sequence Models
- Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things
- Towards better decoding and language model integration in sequence to sequence models
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning
- Challenges in Data-to-Document Generation
- TIPRDC: Task-Independent Privacy-Respecting Data Crowdsourcing Framework for Deep Learning with Anonymized Intermediate Representations
- Learning Algorithms for Active Learning
- Neural Network Distiller: A Python Package For DNN Compression Research
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Text Compression-aided Transformer Encoding
- On the Binding Problem in Artificial Neural Networks
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward Pass
- Differentiable Physics-informed Graph Networks
- Bridging Neural Machine Translation and Bilingual Dictionaries
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Multi-layer Representation Fusion for Neural Machine Translation
- Faster Fuzzing: Reinitialization with Deep Neural Models
- An Unsupervised Autoregressive Model for Speech Representation Learning
- Massively Multilingual Neural Machine Translation
- Neural Information Retrieval: A Literature Review
- Generative Language Modeling for Automated Theorem Proving
- Attention U-Net as a surrogate model for groundwater prediction
- A Selective Overview of Deep Learning
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation
- Mixed Precision Training With 8-bit Floating Point
- Calibration of Encoder Decoder Models for Neural Machine Translation
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
- Local minima in training of neural networks
- DeepRecSys: A System for Optimizing End-To-End At-scale Neural Recommendation Inference
- Robust Neural Machine Translation with Doubly Adversarial Inputs
- End-to-End Speech Translation with Knowledge Distillation
- Scale MLPerf-0.6 models on Google TPU-v3 Pods
- CG-BERT: Conditional Text Generation with BERT for Generalized Few-shot Intent Detection
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- Passport-aware Normalization for Deep Model Protection
- From Zero to Hero: On the Limitations of Zero-Shot Cross-Lingual Transfer with Multilingual Transformers
- Hierarchical Transformers for Multi-Document Summarization
- Controlling the Output Length of Neural Machine Translation
- Neural Text Generation: A Practical Guide
- DRAGNN: A Transition-based Framework for Dynamically Connected Neural Networks
- A Call for Prudent Choice of Subword Merge Operations in Neural Machine Translation
- Progress Notes Classification and Keyword Extraction using Attention-based Deep Learning Models with BERT
- Diagnosing BERT with Retrieval Heuristics
- Placeto: Learning Generalizable Device Placement Algorithms for Distributed Machine Learning
- TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
- Stochastic Mirror Descent on Overparameterized Nonlinear Models: Convergence, Implicit Regularization, and Generalization
- Machine Reading Comprehension: a Literature Review
- DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
- They, Them, Theirs: Rewriting with Gender-Neutral English
- The Impact of Preprocessing on Arabic-English Statistical and Neural Machine Translation
- ST-HOI: A Spatial-Temporal Baseline for Human-Object Interaction Detection in Videos
- CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters
- Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models
- GDP: Generalized Device Placement for Dataflow Graphs
- Online Learning for Neural Machine Translation Post-editing
- Platform for Situated Intelligence
- Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
- CharacterBERT and Self-Teaching for Improving the Robustness of Dense Retrievers on Queries with Typos
- Machine-Translation History and Evolution: Survey for Arabic-English Translations
- Query-Based Abstractive Summarization Using Neural Networks
- KR-BERT: A Small-Scale Korean-Specific Language Model
- Can We Generate Shellcodes via Natural Language? An Empirical Study
- Context in Neural Machine Translation: A Review of Models and Evaluations
- Neural Language Generation: Formulation, Methods, and Evaluation
- Improving Sequence-to-Sequence Learning via Optimal Transport
- A Comparative Study of Transformer-Based Language Models on Extractive Question Answering
- A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data Poisoning
- Direct speech-to-speech translation with a sequence-to-sequence model
- Bridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog
- Deep Learning for Scene Classification: A Survey
- Cold-Start Reinforcement Learning with Softmax Policy Gradient
- An Exploration of Word Embedding Initialization in Deep-Learning Tasks
- BioNetExplorer: Architecture-Space Exploration of Bio-Signal Processing Deep Neural Networks for Wearables
- BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization
- A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
- Mitigating Edge Machine Learning Inference Bottlenecks: An Empirical Study on Accelerating Google Edge Models
- Confidence through Attention
- Does injecting linguistic structure into language models lead to better alignment with brain recordings?
- Improving Readability for Automatic Speech Recognition Transcription
- Scaling Laws for Neural Machine Translation
- Implicit Dimension Identification in User-Generated Text with LSTM Networks
- Automatic Cross-Replica Sharding of Weight Update in Data-Parallel Training
- Wat zei je? Detecting Out-of-Distribution Translations with Variational Transformers
- Unsupervised Neural Machine Translation with SMT as Posterior Regularization
- Contrastive Visual-Linguistic Pretraining
- Federated Learning of N-gram Language Models
- Memory-augmented Neural Machine Translation
- Stochastic Beams and Where to Find Them: The Gumbel-Top-k Trick for Sampling Sequences Without Replacement
- Multi-Channel Auto-Calibration for the Atmospheric Imaging Assembly using Machine Learning
- Learning to Focus: Cascaded Feature Matching Network for Few-shot Image Recognition
- Exploring Benefits of Transfer Learning in Neural Machine Translation
- Competence-based Curriculum Learning for Neural Machine Translation
- Non-Autoregressive Neural Machine Translation with Enhanced Decoder Input
- Adversarial Training with Voronoi Constraints
- Context-Aware Learning for Neural Machine Translation
- Pretrained Language Models for Document-Level Neural Machine Translation
- Bit-Flip Attack: Crushing Neural Network with Progressive Bit Search
- Pre-Trained Models: Past, Present and Future
- Analogue Quantum Simulation: A New Instrument for Scientific Understanding
- Extracting UMLS Concepts from Medical Text Using General and Domain-Specific Deep Learning Models
- Anomaly Detection in Multivariate Non-stationary Time Series for Automatic DBMS Diagnosis
- Neural Machine Translating from Natural Language to SPARQL
- Multi-branch Attentive Transformer
- Multilingual Name Entity Recognition and Intent Classification Employing Deep Learning Architectures
- Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement
- Cross-Lingual Semantic Role Labeling with High-Quality Translated Training Corpus
- SS-Auto: A Single-Shot, Automatic Structured Weight Pruning Framework of DNNs with Ultra-High Efficiency
- A Deep Reinforcement Learning Chatbot (Short Version)
- PAWS: Paraphrase Adversaries from Word Scrambling
- Liputan6: A Large-scale Indonesian Dataset for Text Summarization
- Trainable Greedy Decoding for Neural Machine Translation
- Deep Recurrent Neural Network for Protein Function Prediction from Sequence
- Match-Tensor: a Deep Relevance Model for Search
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
- Multimodal Joint Attribute Prediction and Value Extraction for E-commerce Product
- Accelerating Sparse Deep Neural Networks
- Towards Understanding the Faults of JavaScript-Based Deep Learning Systems
- Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
- VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference
- Natural Question Generation with Reinforcement Learning Based Graph-to-Sequence Model
- Adversarial Attacks and Defense on Texts: A Survey
- Information Aggregation for Multi-Head Attention with Routing-by-Agreement
- PharmMT: A Neural Machine Translation Approach to Simplify Prescription Directions
- Why Do Masked Neural Language Models Still Need Common Sense Knowledge?
- Bayesian Sparsification of Recurrent Neural Networks
- SEAL: Segment-wise Extractive-Abstractive Long-form Text Summarization
- BoostingBERT:Integrating Multi-Class Boosting into BERT for NLP Tasks
- Knowledge Enhanced Contextual Word Representations
- A Correspondence Between Random Neural Networks and Statistical Field Theory
- Tiny Transformers for Environmental Sound Classification at the Edge
- Towards Preemptive Detection of Depression and Anxiety in Twitter
- QAScore -- An Unsupervised Unreferenced Metric for the Question Generation Evaluation
- A Research Agenda: Dynamic Models to Defend Against Correlated Attacks
- One-shot and few-shot learning of word embeddings
- Steering Output Style and Topic in Neural Response Generation
- Simplify-then-Translate: Automatic Preprocessing for Black-Box Machine Translation
- Domain-specific Communication Optimization for Distributed DNN Training
- Tree Transformer: Integrating Tree Structures into Self-Attention
- Toward a Period-Specific Optimized Neural Network for OCR Error Correction of Historical Hebrew Texts
- On the Interpolation of Contextualized Term-based Ranking with BM25 for Query-by-Example Retrieval
- Translating Phrases in Neural Machine Translation
- Personalized Re-ranking for Recommendation
- Low-Precision Batch-Normalized Activations
- Neural network gradient-based learning of black-box function interfaces
- Hard-Coded Gaussian Attention for Neural Machine Translation
- Code-switched inspired losses for generic spoken dialog representations
- SimCLS: A Simple Framework for Contrastive Learning of Abstractive Summarization
- Efficient Attentions for Long Document Summarization
- CREDIT: Coarse-to-Fine Sequence Generation for Dialogue State Tracking
- Incremental Adaptation of NMT for Professional Post-editors: A User Study
- Retrieving Sequential Information for Non-Autoregressive Neural Machine Translation
- Improving Sign Language Translation with Monolingual Data by Sign Back-Translation
- Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
- SMT vs NMT: A Comparison over Hindi & Bengali Simple Sentences
- BERTSurv: BERT-Based Survival Models for Predicting Outcomes of Trauma Patients
- Beam Search with Bidirectional Strategies for Neural Response Generation
- What does Attention in Neural Machine Translation Pay Attention to?
- BERT-DST: Scalable End-to-End Dialogue State Tracking with Bidirectional Encoder Representations from Transformer
- POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-training
- Analysis of Predictive Coding Models for Phonemic Representation Learning in Small Datasets
- Named Entity Recognition in the Legal Domain using a Pointer Generator Network
- Exploring the limits of Concurrency in ML Training on Google TPUs
- Neural Machine Translation: Challenges, Progress and Future
- KAISA: An Adaptive Second-Order Optimizer Framework for Deep Neural Networks
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning
- CS1QA: A Dataset for Assisting Code-based Question Answering in an Introductory Programming Course
- Adversarial Learning for Chinese NER from Crowd Annotations
- Non-Autoregressive Text Generation with Pre-trained Language Models
- Data Augmentation for Low-Resource Named Entity Recognition Using Backtranslation
- Modelling Latent Translations for Cross-Lingual Transfer
- Low-Resource Language Modelling of South African Languages
- Exploring Architectures, Data and Units For Streaming End-to-End Speech Recognition with RNN-Transducer
- AMR Parsing as Sequence-to-Graph Transduction
- Compositional Generalization for Primitive Substitutions
- Improving Biomedical Pretrained Language Models with Knowledge
- Efficient and Generic 1D Dilated Convolution Layer for Deep Learning
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Literature review on vulnerability detection using NLP technology
- Evaluating Transfer Learning for Simplifying GitHub READMEs
- Literature Retrieval for Precision Medicine with Neural Matching and Faceted Summarization
- tf.data: A Machine Learning Data Processing Framework
- HanoiT: Enhancing Context-aware Translation via Selective Context
- Language as a matrix product state
- Scalable Transformers for Neural Machine Translation
- StructuralLM: Structural Pre-training for Form Understanding
- XtremeDistilTransformers: Task Transfer for Task-agnostic Distillation
- Introducing Aspects of Creativity in Automatic Poetry Generation
- Speech-language Pre-training for End-to-end Spoken Language Understanding
- Do Transformer Attention Heads Provide Transparency in Abstractive Summarization?
- An Empirical Study of Efficient ASR Rescoring with Transformers
- Multi-object Tracking via End-to-end Tracklet Searching and Ranking
- Biomedical Entity Representations with Synonym Marginalization
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
- A Study of Multilingual Neural Machine Translation
- Cross-lingual Pre-training Based Transfer for Zero-shot Neural Machine Translation
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Character-based NMT with Transformer
- BaPipe: Exploration of Balanced Pipeline Parallelism for DNN Training
- Large-scale Pretraining for Neural Machine Translation with Tens of Billions of Sentence Pairs
- A Latent Morphology Model for Open-Vocabulary Neural Machine Translation
- Tell Me Why You Feel That Way: Processing Compositional Dependency for Tree-LSTM Aspect Sentiment Triplet Extraction (TASTE)
- A Computational Model of Commonsense Moral Decision Making
- ERNIE at SemEval-2020 Task 10: Learning Word Emphasis Selection by Pre-trained Language Model
- SAFER: A Structure-free Approach for Certified Robustness to Adversarial Word Substitutions
- Norm-Based Curriculum Learning for Neural Machine Translation
- Synchronous Bidirectional Inference for Neural Sequence Generation
- High-Performance Deep Learning via a Single Building Block
- A Comparison of Approaches to Document-level Machine Translation
- End-to-End Speech Recognition: A review for the French Language
- Pre-training Text Representations as Meta Learning
- HateProof: Are Hateful Meme Detection Systems really Robust?
- Importance-Aware Learning for Neural Headline Editing
- iCapsNets: Towards Interpretable Capsule Networks for Text Classification
- Neural Chinese Word Segmentation as Sequence to Sequence Translation
- CNN Is All You Need
- LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
- On the quantization of recurrent neural networks
- Multi-node Bert-pretraining: Cost-efficient Approach
- Mono vs Multilingual Transformer-based Models: a Comparison across Several Language Tasks
- UM-IU@LING at SemEval-2019 Task 6: Identifying Offensive Tweets Using BERT and SVMs
- ReWE: Regressing Word Embeddings for Regularization of Neural Machine Translation Systems
- Modelling Sequential Music Track Skips using a Multi-RNN Approach
- Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak Supervision (Short Version)
- GNN-XML: Graph Neural Networks for Extreme Multi-label Text Classification
- Class LM and word mapping for contextual biasing in End-to-End ASR
- Intelligent Video Editing: Incorporating Modern Talking Face Generation Algorithms in a Video Editor
- BERT as a Teacher: Contextual Embeddings for Sequence-Level Reward
- Fast Sequence Generation with Multi-Agent Reinforcement Learning
- Adversarial Learning of Privacy-Preserving and Task-Oriented Representations
- SkillSpan: Hard and Soft Skill Extraction from English Job Postings
- Sharing Attention Weights for Fast Transformer
- Multiscale Collaborative Deep Models for Neural Machine Translation
- Jointly Optimizing Diversity and Relevance in Neural Response Generation
- KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning
- Grid Search Hyperparameter Benchmarking of BERT, ALBERT, and LongFormer on DuoRC
- AdvAug: Robust Adversarial Augmentation for Neural Machine Translation
- Towards Neural Machine Translation with Partially Aligned Corpora
- Improving the Performance of Online Neural Transducer Models
- Streaming Object Detection for 3-D Point Clouds
- Variable Name Recovery in Decompiled Binary Code using Constrained Masked Language Modeling
- Horizontally Fused Training Array: An Effective Hardware Utilization Squeezer for Training Novel Deep Learning Models
- On Accelerating Distributed Convex Optimizations
- Students Need More Attention: BERT-based AttentionModel for Small Data with Application to AutomaticPatient Message Triage
- SuperChat: Dialogue Generation by Transfer Learning from Vision to Language using Two-dimensional Word Embedding and Pretrained ImageNet CNN Models
- Reward Optimization for Neural Machine Translation with Learned Metrics
- Psycholinguistic Tripartite Graph Network for Personality Detection
- Superbloom: Bloom filter meets Transformer
- Text Analysis in Adversarial Settings: Does Deception Leave a Stylistic Trace?
- Analysis of the Evolution of Advanced Transformer-Based Language Models: Experiments on Opinion Mining
- Multi-Head Multi-Layer Attention to Deep Language Representations for Grammatical Error Detection
- Named entity recognition in chemical patents using ensemble of contextual language models
- FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval
- Understanding How Encoder-Decoder Architectures Attend
- StrucTexT: Structured Text Understanding with Multi-Modal Transformers
- Autocorrect in the Process of Translation -- Multi-task Learning Improves Dialogue Machine Translation
- Recent Trends in the Use of Deep Learning Models for Grammar Error Handling
- Attentional Speech Recognition Models Misbehave on Out-of-domain Utterances
- Teaching Machines to Converse
- Composed Variational Natural Language Generation for Few-shot Intents
- Deep Ensembles on a Fixed Memory Budget: One Wide Network or Several Thinner Ones?
- Multi-Task Learning for Conversational Question Answering over a Large-Scale Knowledge Base
- Reinforced Iterative Knowledge Distillation for Cross-Lingual Named Entity Recognition
- Multimodal Learning for Hateful Memes Detection
- Towards Unsupervised Language Understanding and Generation by Joint Dual Learning
- Multimodal Dialogue State Tracking By QA Approach with Data Augmentation
- Cutting-off Redundant Repeating Generations for Neural Abstractive Summarization
- UniTrans: Unifying Model Transfer and Data Transfer for Cross-Lingual Named Entity Recognition with Unlabeled Data
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting
- MARMOT: A Deep Learning Framework for Constructing Multimodal Representations for Vision-and-Language Tasks
- AI-Powered Social Bots
- How can AI Automate End-to-End Data Science?
- Why Neural Machine Translation Prefers Empty Outputs
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- Automatic Post-Editing for Machine Translation
- Interactive-Predictive Neural Machine Translation through Reinforcement and Imitation
- Guiding High-Performance SAT Solvers with Unsat-Core Predictions
- PiSLTRc: Position-informed Sign Language Transformer with Content-aware Convolution
- Low-Shot Classification: A Comparison of Classical and Deep Transfer Machine Learning Approaches
- Translating Translationese: A Two-Step Approach to Unsupervised Machine Translation
- Word-based Domain Adaptation for Neural Machine Translation
- TauRieL: Targeting Traveling Salesman Problem with a deep reinforcement learning inspired architecture
- Text Generation with Exemplar-based Adaptive Decoding
- A Convolutional Neural Network for Language-Agnostic Source Code Summarization
- Learning Efficient Lexically-Constrained Neural Machine Translation with External Memory
- Hierarchical Transformer Network for Utterance-level Emotion Recognition
- Learning from Learning Machines: Optimisation, Rules, and Social Norms
- Seq2Seq RNN based Gait Anomaly Detection from Smartphone Acquired Multimodal Motion Data
- Robust Reading Comprehension with Linguistic Constraints via Posterior Regularization
- Open Domain Web Keyphrase Extraction Beyond Language Modeling
- Positional Attention-based Frame Identification with BERT: A Deep Learning Approach to Target Disambiguation and Semantic Frame Selection
- Attention Forcing for Sequence-to-sequence Model Training
- Improving Pre-Trained Multilingual Models with Vocabulary Expansion
- Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals
- Rethinking Dialogue State Tracking with Reasoning
- Capsule-Transformer for Neural Machine Translation
- A Practical Guide to Studying Emergent Communication through Grounded Language Games
- Learning from Imperfect Annotations
- AR: Auto-Repair the Synthetic Data for Neural Machine Translation
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set Synthesis
- HufuNet: Embedding the Left Piece as Watermark and Keeping the Right Piece for Ownership Verification in Deep Neural Networks
- Enriching Non-Autoregressive Transformer with Syntactic and SemanticStructures for Neural Machine Translation
- Skin disease diagnosis with deep learning: a review
- CapWAP: Captioning with a Purpose
- Fast Interleaved Bidirectional Sequence Generation
- Single-Read Reconstruction for DNA Data Storage Using Transformers
- Regressive Ensemble for Machine Translation Quality Evaluation
- IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization
- Log-based Anomaly Detection Without Log Parsing
- Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and Conformers
- AdvPicker: Effectively Leveraging Unlabeled Data via Adversarial Discriminator for Cross-Lingual NER
- SGD-QA: Fast Schema-Guided Dialogue State Tracking for Unseen Services
- Can NMT Understand Me? Towards Perturbation-based Evaluation of NMT Models for Code Generation
- LOCAL: Low-Complex Mapping Algorithm for Spatial DNN Accelerators
- On-the-Fly Syntax Highlighting using Neural Networks
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- Bridging Subword Gaps in Pretrain-Finetune Paradigm for Natural Language Generation
- Contextualized Rewriting for Text Summarization
- What's in a Name? -- Gender Classification of Names with Character Based Machine Learning Models
- Time to Take Emoji Seriously: They Vastly Improve Casual Conversational Models
- Improving Human Text Comprehension through Semi-Markov CRF-based Neural Section Title Generation
- BanditRank: Learning to Rank Using Contextual Bandits
- UzBERT: pretraining a BERT model for Uzbek
- Multi-Instance Learning for End-to-End Knowledge Base Question Answering
- Efficient Contextual Representation Learning Without Softmax Layer
- Revisiting Language Encoding in Learning Multilingual Representations
- Is In-hospital Meta-information Useful for Abstractive Discharge Summary Generation?
- Zero-shot Dependency Parsing with Pre-trained Multilingual Sentence Representations
- A Survey of Techniques All Classifiers Can Learn from Deep Networks: Models, Optimizations, and Regularization
- Recurrent Neural Network-based Model for Accelerated Trajectory Analysis in AIMD Simulations
- Semi-Supervised Learning with Data Augmentation for End-to-End ASR
- Noisy Text Data: Achilles' Heel of popular transformer based NLP models
- MANet: Multimodal Attention Network based Point- View fusion for 3D Shape Recognition
- Transferring Monolingual Model to Low-Resource Language: The Case of Tigrinya
- PERL: Pivot-based Domain Adaptation for Pre-trained Deep Contextualized Embedding Models
- An Empirical Accuracy Law for Sequential Machine Translation: the Case of Google Translate
- Speeding up Deep Model Training by Sharing Weights and Then Unsharing
- Toward Computation and Memory Efficient Neural Network Acoustic Models with Binary Weights and Activations
- Contextual embedding and model weighting by fusing domain knowledge on Biomedical Question Answering
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
- Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
- AdsGNN: Behavior-Graph Augmented Relevance Modeling in Sponsored Search
- The Mathematical Foundations of Manifold Learning
- Assessing Demographic Bias in Named Entity Recognition
- Explicit Sentence Compression for Neural Machine Translation
- Neural Text Generation with Artificial Negative Examples
- Sinhala-English Parallel Word Dictionary Dataset
- Stage-based Hyper-parameter Optimization for Deep Learning
- SubCharacter Chinese-English Neural Machine Translation with Wubi encoding
- Alleviating the Burden of Labeling: Sentence Generation by Attention Branch Encoder-Decoder Network
- EPNAS: Efficient Progressive Neural Architecture Search
- Improving Bidirectional Decoding with Dynamic Target Semantics in Neural Machine Translation
- Event Detection as Question Answering with Entity Information
- Conversational Question Answering: A Survey
- "The Boating Store Had Its Best Sail Ever": Pronunciation-attentive Contextualized Pun Recognition
- Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
- Vulnerability Under Adversarial Machine Learning: Bias or Variance?
- Architecture for a multilingual Wikipedia
- Understanding Learning Dynamics for Neural Machine Translation
- Addressing the Vulnerability of NMT in Input Perturbations
- Patient Cohort Retrieval using Transformer Language Models
- Margin-Based Regularization and Selective Sampling in Deep Neural Networks
- Understanding Transformers for Bot Detection in Twitter
- AI Powered Compiler Techniques for DL Code Optimization
- Sentiment-based Candidate Selection for NMT
- Lessons Learned from Applying off-the-shelf BERT: There is no Silver Bullet
- Being-ahead: Benchmarking and Exploring Accelerators for Hardware-Efficient AI Deployment
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- Attention Forcing for Machine Translation
- Evaluating Neural Word Embeddings for Sanskrit
- PAIR: Planning and Iterative Refinement in Pre-trained Transformers for Long Text Generation
- Iterative Domain-Repaired Back-Translation
- On the Importance of Subword Information for Morphological Tasks in Truly Low-Resource Languages
- Classifying the Cosmic-Ray Proton and Light Groups on the LHAASO-KM2A Experiment with the Graph Neural Network
- Using Interlinear Glosses as Pivot in Low-Resource Multilingual Machine Translation
- Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge
- Human-centric Metric for Accelerating Pathology Reports Annotation
- Learning Light-Weight Translation Models from Deep Transformer
- Neural Machine Translation with Explicit Phrase Alignment
- A Feasible Framework for Arbitrary-Shaped Scene Text Recognition
- Traditional IR rivals neural models on the MS MARCO Document Ranking Leaderboard
- Multilingual Word Embeddings using Multigraphs
- Collaborative Training of GANs in Continuous and Discrete Spaces for Text Generation
- Investigation of BERT Model on Biomedical Relation Extraction Based on Revised Fine-tuning Mechanism
- On the Application of Data-Driven Deep Neural Networks in Linear and Nonlinear Structural Dynamics
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition
- Improving Neural Machine Translation by Bidirectional Training
- Enhanced Neural Machine Translation by Learning from Draft
- High-performance stock index trading: making effective use of a deep LSTM neural network
- A spelling correction model for end-to-end speech recognition
- Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning
- UIBert: Learning Generic Multimodal Representations for UI Understanding
- Sequence Model with Self-Adaptive Sliding Window for Efficient Spoken Document Segmentation
- Syntax-Enhanced Neural Machine Translation with Syntax-Aware Word Representations
- Synchronous Bidirectional Neural Machine Translation
- Residual Error: a New Performance Measure for Adversarial Robustness
- Language Tags Matter for Zero-Shot Neural Machine Translation
- GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
- Adversarial Sub-sequence for Text Generation
- Sequence-Level Training for Non-Autoregressive Neural Machine Translation
- Modeling Past and Future for Neural Machine Translation
- Using Linguistic Typology to Enrich Multilingual Lexicons: the Case of Lexical Gaps in Kinship
- Semantic Parsing with Dual Learning
- GraphCode2Vec: Generic Code Embedding via Lexical and Program Dependence Analyses
- Hippo: Taming Hyper-parameter Optimization of Deep Learning with Stage Trees
- MOOCRep: A Unified Pre-trained Embedding of MOOC Entities
- FastSeq: Make Sequence Generation Faster
- Word, Subword or Character? An Empirical Study of Granularity in Chinese-English NMT
- Distributional Discrepancy: A Metric for Unconditional Text Generation
- SLAM-Inspired Simultaneous Contextualization and Interpreting for Incremental Conversation Sentences
- DynamicEmbedding: Extending TensorFlow for Colossal-Scale Applications
- Exploring Pre-training with Alignments for RNN Transducer based End-to-End Speech Recognition
- Jointly Encoding Word Confusion Network and Dialogue Context with BERT for Spoken Language Understanding
- BERT-JAM: Boosting BERT-Enhanced Neural Machine Translation with Joint Attention
- Learning a Reinforced Agent for Flexible Exposure Bracketing Selection
- DCoM: A Deep Column Mapper for Semantic Data Type Detection
- Depth-Adaptive Neural Networks from the Optimal Control viewpoint
- Sketching Transformed Matrices with Applications to Natural Language Processing
- Largest Eigenvalues of the Conjugate Kernel of Single-Layered Neural Networks
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- Investigating Label Bias in Beam Search for Open-ended Text Generation
- Amharic-Arabic Neural Machine Translation
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- Interpretable Self-supervised Multi-task Learning for COVID-19 Information Retrieval and Extraction
- A Lightweight Recurrent Network for Sequence Modeling
- Guiding Teacher Forcing with Seer Forcing for Neural Machine Translation
- SQWA: Stochastic Quantized Weight Averaging for Improving the Generalization Capability of Low-Precision Deep Neural Networks
- To Beam Or Not To Beam: That is a Question of Cooperation for Language GANs
- Aligning the Pretraining and Finetuning Objectives of Language Models
- BERT might be Overkill: A Tiny but Effective Biomedical Entity Linker based on Residual Convolutional Neural Networks
- Lattice Rescoring Strategies for Long Short Term Memory Language Models in Speech Recognition
- MergeDistill: Merging Pre-trained Language Models using Distillation
- Dynamically Composing Domain-Data Selection with Clean-Data Selection by "Co-Curricular Learning" for Neural Machine Translation
- Defending Democracy: Using Deep Learning to Identify and Prevent Misinformation
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- A Semi-Supervised Approach for Low-Resourced Text Generation
- Convolutional Attention-based Seq2Seq Neural Network for End-to-End ASR
- Automatic Dialogic Instruction Detection for K-12 Online One-on-one Classes
- Understanding Natural Language Instructions for Fetching Daily Objects Using GAN-Based Multimodal Target-Source Classification
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- Integration of TensorFlow based Acoustic Model with Kaldi WFST Decoder
- A Multilingual Modeling Method for Span-Extraction Reading Comprehension
- GWLAN: General Word-Level AutocompletioN for Computer-Aided Translation
- Search Spaces for Neural Model Training
- Fast and Accurate Deep Bidirectional Language Representations for Unsupervised Learning
- Goal-driven text descriptions for images
- Deceptive Deletions for Protecting Withdrawn Posts on Social Platforms
- Adma: A Flexible Loss Function for Neural Networks
- Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks
- RLTIR: Activity-based Interactive Person Identification based on Reinforcement Learning Tree
- Variational Knowledge Distillation for Disease Classification in Chest X-Rays
- Sensor Transformation Attention Networks
- Smoothing and Shrinking the Sparse Seq2Seq Search Space
- Depth Growing for Neural Machine Translation
- Improved Speech Representations with Multi-Target Autoregressive Predictive Coding
- Modeling Coreference Relations in Visual Dialog
- Towards Fully Automated Manga Translation
- Artificial Intelligence : from Research to Application ; the Upper-Rhine Artificial Intelligence Symposium (UR-AI 2019)
- Experiments with Rich Regime Training for Deep Learning
- Bayesian Sparsification of Gated Recurrent Neural Networks
- Language Transfer for Early Warning of Epidemics from Social Media
- Breaking the Memory Wall for AI Chip with a New Dimension
- Learning Architectures from an Extended Search Space for Language Modeling
- Correction of Faulty Background Knowledge based on Condition Aware and Revise Transformer for Question Answering
- Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based Games
- Crowdsourcing Parallel Corpus for English-Oromo Neural Machine Translation using Community Engagement Platform
- An Intrinsic Nearest Neighbor Analysis of Neural Machine Translation Architectures
- Vanishing Nodes: Another Phenomenon That Makes Training Deep Neural Networks Difficult
- ML Based Lineage in Databases
- Understand customer reviews with less data and in short time: pretrained language representation and active learning
- An Optimized and Energy-Efficient Parallel Implementation of Non-Iteratively Trained Recurrent Neural Networks
- A Sketch-Based Neural Model for Generating Commit Messages from Diffs
- BERT-based Chinese Text Classification for Emergency Domain with a Novel Loss Function
- Knowledge Efficient Deep Learning for Natural Language Processing
- Indirectly Supervised English Sentence Break Prediction Using Paragraph Break Probability Estimates
- ActBERT: Learning Global-Local Video-Text Representations
- A Comprehensive Study on Temporal Modeling for Online Action Detection
- DeepTitle -- Leveraging BERT to generate Search Engine Optimized Headlines
- Multilingual Neural RST Discourse Parsing
- On the Copying Behaviors of Pre-Training for Neural Machine Translation
- Phases of learning dynamics in artificial neural networks: with or without mislabeled data
- XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages
- Cross-Modality Relevance for Reasoning on Language and Vision
- Modeling Diagnostic Label Correlation for Automatic ICD Coding
- All You Can Embed: Natural Language based Vehicle Retrieval with Spatio-Temporal Transformers
- Learning to Detect Unacceptable Machine Translations for Downstream Tasks
- Multi-Task Learning and Adapted Knowledge Models for Emotion-Cause Extraction
- Algorithm to Compilation Co-design: An Integrated View of Neural Network Sparsity
- Dynamic Programming Encoding for Subword Segmentation in Neural Machine Translation
- Morphological Skip-Gram: Using morphological knowledge to improve word representation
- MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural Networks
- TAG : Type Auxiliary Guiding for Code Comment Generation
- Multi-modal Transformer for Video Retrieval
- Improving Truthfulness of Headline Generation
- Keeping Notes: Conditional Natural Language Generation with a Scratchpad Mechanism
- Conversational Response Re-ranking Based on Event Causality and Role Factored Tensor Event Embedding
- The University of Sydney's Machine Translation System for WMT19
- Neural Machine Translation for Multilingual Grapheme-to-Phoneme Conversion
- NameRec*: Highly Accurate and Fine-grained Person Name Recognition
- Learning to Reformulate the Queries on the WEB
- An empirical analysis of phrase-based and neural machine translation
- Data Augmentation for Abstractive Query-Focused Multi-Document Summarization
- How LSTM Encodes Syntax: Exploring Context Vectors and Semi-Quantization on Natural Text
- Adapting MARBERT for Improved Arabic Dialect Identification: Submission to the NADI 2021 Shared Task
- Translation, Sentiment and Voices: A Computational Model to Translate and Analyse Voices from Real-Time Video Calling
- Understanding and Enhancing the Use of Context for Machine Translation
- BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge
- Efficient transfer learning for NLP with ELECTRA
- Customizing Contextualized Language Models forLegal Document Reviews
- Single-Queue Decoding for Neural Machine Translation
- On Multilingual Training of Neural Dependency Parsers
- Machine Translation in Pronunciation Space
- Deep Collective Learning: Learning Optimal Inputs and Weights Jointly in Deep Neural Networks
- Training Neural Machine Translation (NMT) Models using Tensor Train Decomposition on TensorFlow (T3F)
- Cross-Modal Alignment with Mixture Experts Neural Network for Intral-City Retail Recommendation
- The Helsinki Neural Machine Translation System
- Duluth at SemEval-2020 Task 7: Using Surprise as a Key to Unlock Humorous Headlines
- Learning to Copy for Automatic Post-Editing
- Stylized Text Generation Using Wasserstein Autoencoders with a Mixture of Gaussian Prior
- SYSTRAN Purely Neural MT Engines for WMT2017
- A crossover code for high-dimensional composition
- The Agnostic Structure of Data Science Methods
- Langevin Cooling for Domain Translation
- The Detection of Distributional Discrepancy for Text Generation
- AGenT Zero: Zero-shot Automatic Multiple-Choice Question Generation for Skill Assessments
- NTT's Machine Translation Systems for WMT19 Robustness Task
- Robust, Deep, and Reinforcement Learning for Management of Communication and Power Networks
- Iterative Dual Domain Adaptation for Neural Machine Translation
- Artificial stochastic neural network on the base of double quantum wells
- Shareable Representations for Search Query Understanding
- Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression
- Dynamic Curriculum Learning for Low-Resource Neural Machine Translation
- Accelerating Distributed SGD for Linear Regression using Iterative Pre-Conditioning
- LfEdNet: A Task-based Day-ahead Load Forecasting Model for Stochastic Economic Dispatch
- Empirical Evaluation of Deep Learning Model Compression Techniques on the WaveNet Vocoder
- Bi-Decoder Augmented Network for Neural Machine Translation
- DORB: Dynamically Optimizing Multiple Rewards with Bandits
- Multilingual AMR-to-Text Generation
- Pretrained Transformers for Simple Question Answering over Knowledge Graphs
- Layer-Wise Multi-View Learning for Neural Machine Translation
- PolyScientist: Automatic Loop Transformations Combined with Microkernels for Optimization of Deep Learning Primitives
- Learning to Generate Multiple Style Transfer Outputs for an Input Sentence
- Unsupervised Learning Layers for Video Analysis
- Residual Continual Learning
- Modeling Coverage for Non-Autoregressive Neural Machine Translation
- Optimizing AD Pruning of Sponsored Search with Reinforcement Learning
- UPB at SemEval-2020 Task 12: Multilingual Offensive Language Detection on Social Media by Fine-tuning a Variety of BERT-based Models
- Long Short-Term Memory Neuron Equalizer
- Reducing Data Motion to Accelerate the Training of Deep Neural Networks
- Phantom: A High-Performance Computational Core for Sparse Convolutional Neural Networks
- Image to Language Understanding: Captioning approach
- Explaining Documents' Relevance to Search Queries
- PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation
- Guider l'attention dans les modeles de sequence a sequence pour la prediction des actes de dialogue
- An Empirical Investigation of Multi-bridge Multilingual NMT models
- Subspace Approximation for Approximate Nearest Neighbor Search in NLP
- Neural Sequence Model Training via -divergence Minimization
- Measuring prominence of scientific work in online news as a proxy for impact
- :An Unbiased Stratified Statistic and a Fast Gradient Optimization Algorithm Based on It
- Integrating Categorical Features in End-to-End ASR
- Improving Zero-shot Multilingual Neural Machine Translation for Low-Resource Languages
- Effective Use of Graph Convolution Network and Contextual Sub-Tree forCommodity News Event Extraction
- Sentence-wise Smooth Regularization for Sequence to Sequence Learning
- Enhancing Clinical Information Extraction with Transferred Contextual Embeddings
- Task-adaptive Pre-training of Language Models with Word Embedding Regularization
- Open-endedness in AI systems, cellular evolution and intellectual discussions
- A Massively Multilingual Analysis of Cross-linguality in Shared Embedding Space
- Application of Video-to-Video Translation Networks to Computational Fluid Dynamics
- Sequential Modelling with Applications to Music Recommendation, Fact-Checking, and Speed Reading
- Style Pooling: Automatic Text Style Obfuscation for Improved Classification Fairness
- An Evaluation Dataset and Strategy for Building Robust Multi-turn Response Selection Model
- ARMAN: Pre-training with Semantically Selecting and Reordering of Sentences for Persian Abstractive Summarization
- DeepZensols: Deep Natural Language Processing Framework
- An Emergency Medical Services Clinical Audit System driven by Named Entity Recognition from Deep Learning
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- Can the Transformer Be Used as a Drop-in Replacement for RNNs in Text-Generating GANs?
- Recurrent multiple shared layers in Depth for Neural Machine Translation
- Deploying a BERT-based Query-Title Relevance Classifier in a Production System: a View from the Trenches
- An Empirical Study of Mini-Batch Creation Strategies for Neural Machine Translation
- Identifying and Exploiting Structures for Reliable Deep Learning
- Dialogue Summarization with Supporting Utterance Flow Modeling and Fact Regularization
- Using Structured Input and Modularity for Improved Learning
- An End-to-End Approach to Automatic Speech Assessment for Cantonese-speaking People with Aphasia
- The Cross-Lingual Arabic Information REtrieval (CLAIRE) System
- Learning to Decipher Hate Symbols
- Using CollGram to Compare Formulaic Language in Human and Neural Machine Translation
- Corpora Generation for Grammatical Error Correction
- Residual Tree Aggregation of Layers for Neural Machine Translation
- A Topic Guided Pointer-Generator Model for Generating Natural Language Code Summaries
- Deep Hierarchical Classification for Category Prediction in E-commerce System
- Leveraging Personal Navigation Assistant Systems Using Automated Social Media Traffic Reporting
- Power Law Graph Transformer for Machine Translation and Representation Learning