Convolutional Sequence to Sequence Learning
arXiv:1705.03122
Abstract
The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent models, computations over all elements can be fully parallelized during training and optimization is easier since the number of non-linearities is fixed and independent of the input length. Our use of gated linear units eases gradient propagation and we equip each decoder layer with a separate attention module. We outperform the accuracy of the deep LSTM setup of Wu et al. (2016) on both WMT'14 English-German and WMT'14 English-French translation at an order of magnitude faster speed, both on GPU and CPU.
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Quasi-Recurrent Neural Networks
- Encoding Source Language with Convolutional Neural Network for Machine Translation
- Vocabulary Selection Strategies for Neural Machine Translation
- Neural Machine Translation with Recurrent Attention Modeling
- Cutting-off Redundant Repeating Generations for Neural Abstractive Summarization
Cited by in corpus (521)
- BERTScore: Evaluating Text Generation with BERT
- Pre-trained Models for Natural Language Processing: A Survey
- Recent Trends in Deep Learning Based Natural Language Processing
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Deep Learning on Traffic Prediction: Methods, Analysis and Future Directions
- FastSpeech: Fast, Robust and Controllable Text to Speech
- Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting
- Non-Autoregressive Neural Machine Translation
- Generating Long Sequences with Sparse Transformers
- Beyond English-Centric Multilingual Machine Translation
- QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
- Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
- Deep Learning Based Text Classification: A Comprehensive Review
- Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- The Natural Language Decathlon: Multitask Learning as Question Answering
- Improving Graph Neural Network Expressivity via Subgraph Isomorphism Counting
- Hybrid Quantum-Classical Convolutional Neural Networks
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
- Big Bird: Transformers for Longer Sequences
- Spatial-Temporal Transformer Networks for Traffic Flow Forecasting
- Non-Local Recurrent Network for Image Restoration
- Graph-Based Deep Learning for Medical Diagnosis and Analysis: Past, Present and Future
- GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
- GaAN: Gated Attention Networks for Learning on Large and Spatiotemporal Graphs
- A Survey of Knowledge-Enhanced Text Generation
- Depthwise Separable Convolutions for Neural Machine Translation
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
- Stand-Alone Self-Attention in Vision Models
- MASTER: Multi-Aspect Non-local Network for Scene Text Recognition
- Natural Language Processing Advancements By Deep Learning: A Survey
- A Survey of Domain Adaptation for Neural Machine Translation
- On Extended Long Short-term Memory and Dependent Bidirectional Recurrent Neural Network
- Deformable ConvNets v2: More Deformable, Better Results
- ST-GRAT: A Novel Spatio-temporal Graph Attention Network for Accurately Forecasting Dynamically Changing Road Speed
- Language Modeling with Deep Transformers
- Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
- Fast Decoding in Sequence Models using Discrete Latent Variables
- An Intriguing Failing of Convolutional Neural Networks and the CoordConv Solution
- Reservoir Computing with Random Skyrmion Textures
- Deep Residual Correction Network for Partial Domain Adaptation
- Graph2Seq: Graph to Sequence Learning with Attention-based Neural Networks
- Non-local Neural Networks
- Weighted Transformer Network for Machine Translation
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- Perceiver: General Perception with Iterative Attention
- Transforming Question Answering Datasets Into Natural Language Inference Datasets
- STConvS2S: Spatiotemporal Convolutional Sequence to Sequence Network for Weather Forecasting
- Recent Advances in Deep Learning: An Overview
- Emotion Detection on TV Show Transcripts with Sequence-based Convolutional Neural Networks
- Privacy-preserving Artificial Intelligence Techniques in Biomedicine
- Scene Text Detection and Recognition: The Deep Learning Era
- Understanding Back-Translation at Scale
- Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning
- Reaching Human-level Performance in Automatic Grammatical Error Correction: An Empirical Study
- Trellis Networks for Sequence Modeling
- Stable Recurrent Models
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- A Comparison of Transformer and Recurrent Neural Networks on Multilingual Neural Machine Translation
- The Missing Ingredient in Zero-Shot Neural Machine Translation
- It's Morphin' Time! Combating Linguistic Discrimination with Inflectional Perturbations
- Paradigm Shift in Natural Language Processing
- Quasi-hyperbolic momentum and Adam for deep learning
- Controllable Neural Story Plot Generation via Reward Shaping
- Why gradient clipping accelerates training: A theoretical justification for adaptivity
- Fully Convolutional Speech Recognition
- OpenNMT: Neural Machine Translation Toolkit
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling
- Taskmaster-1: Toward a Realistic and Diverse Dialog Dataset
- On Self Modulation for Generative Adversarial Networks
- A Survey of Automatic Generation of Source Code Comments: Algorithms and Techniques
- Very Deep Transformers for Neural Machine Translation
- NSML: A Machine Learning Platform That Enables You to Focus on Your Models
- Graph Neural Networks for Natural Language Processing: A Survey
- CNN+CNN: Convolutional Decoders for Image Captioning
- DeepTrack: Lightweight Deep Learning for Vehicle Path Prediction in Highways
- Neural Abstractive Text Summarization with Sequence-to-Sequence Models
- Spatial Broadcast Decoder: A Simple Architecture for Learning Disentangled Representations in VAEs
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- Relation Networks for Object Detection
- Non-Autoregressive Machine Translation with Disentangled Context Transformer
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
- Encoding word order in complex embeddings
- ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
- Unsupervised Time Series Outlier Detection with Diversity-Driven Convolutional Ensembles -- Extended Version
- Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification
- Abstractive Summarization of Reddit Posts with Multi-level Memory Networks
- Pyramid Feature Attention Network for Saliency detection
- Pervasive Attention: 2D Convolutional Neural Networks for Sequence-to-Sequence Prediction
- Topic-Guided Variational Autoencoders for Text Generation
- Txt2Img-MHN: Remote Sensing Image Generation from Text Using Modern Hopfield Networks
- Towards Accurate Scene Text Recognition with Semantic Reasoning Networks
- Reading Scene Text with Attention Convolutional Sequence Modeling
- Analyzing Uncertainty in Neural Machine Translation
- Next Item Recommendation with Self-Attention
- Deep Learning for Plasma Tomography and Disruption Prediction from Bolometer Data
- Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning
- Double Embeddings and CNN-based Sequence Labeling for Aspect Extraction
- Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification
- Guided Generation of Cause and Effect
- Forecast Network-Wide Traffic States for Multiple Steps Ahead: A Deep Learning Approach Considering Dynamic Non-Local Spatial Correlation and Non-Stationary Temporal Dependency
- Time2Vec: Learning a Vector Representation of Time
- Text Compression-aided Transformer Encoding
- Emergent Translation in Multi-Agent Communication
- Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
- DeLighT: Deep and Light-weight Transformer
- Universal Neural Machine Translation for Extremely Low Resource Languages
- 6GCVAE: Gated Convolutional Variational Autoencoder for IPv6 Target Generation
- Joint Embedding of Words and Labels for Text Classification
- TNCR: Table Net Detection and Classification Dataset
- Data Diversification: A Simple Strategy For Neural Machine Translation
- Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation
- Mixed-Precision Training for NLP and Speech Recognition with OpenSeq2Seq
- TLSAN: Time-aware Long- and Short-term Attention Network for Next-item Recommendation
- Towards Binary-Valued Gates for Robust LSTM Training
- Unifying Human and Statistical Evaluation for Natural Language Generation
- Dynamic Self-Attention : Computing Attention over Words Dynamically for Sentence Embedding
- Deep Learning Based Chatbot Models
- Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions
- A Focus on Neural Machine Translation for African Languages
- Comparative evaluation of CNN architectures for Image Caption Generation
- Abstractive and Extractive Text Summarization using Document Context Vector and Recurrent Neural Networks
- Multilingual NMT with a language-independent attention bridge
- Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
- Letter-Based Speech Recognition with Gated ConvNets
- Energy-Based Reranking: Improving Neural Machine Translation Using Energy-Based Models
- MetaMT,a MetaLearning Method Leveraging Multiple Domain Data for Low Resource Machine Translation
- End-to-End Adversarial Text-to-Speech
- Are Pre-trained Convolutions Better than Pre-trained Transformers?
- Understanding Knowledge Distillation in Non-autoregressive Machine Translation
- Improving Neural Machine Translation Robustness via Data Augmentation: Beyond Back Translation
- Character-Based Handwritten Text Transcription with Attention Networks
- Data-driven Summarization of Scientific Articles
- Identifying and Controlling Important Neurons in Neural Machine Translation
- Pedestrian Attribute Recognition: A Survey
- Convolutional Sequence to Sequence Model for Human Dynamics
- DivGraphPointer: A Graph Pointer Network for Extracting Diverse Keyphrases
- Relation Distillation Networks for Video Object Detection
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Who Needs Words? Lexicon-Free Speech Recognition
- Meta-Learning for Low-Resource Neural Machine Translation
- Describe and Attend to Track: Learning Natural Language guided Structural Representation and Visual Attention for Object Tracking
- Towards a Robust Deep Neural Network in Texts: A Survey
- scb-mt-en-th-2020: A Large English-Thai Parallel Corpus
- Controllable Abstractive Summarization
- An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
- A Reinforced Topic-Aware Convolutional Sequence-to-Sequence Model for Abstractive Text Summarization
- Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling
- Learning to Segment Inputs for NMT Favors Character-Level Processing
- On Extractive and Abstractive Neural Document Summarization with Transformer Language Models
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling
- Parallelizing Linear Recurrent Neural Nets Over Sequence Length
- ProGraML: Graph-based Deep Learning for Program Optimization and Analysis
- Exploiting Deep Representations for Neural Machine Translation
- Improving Grammatical Error Correction via Pre-Training a Copy-Augmented Architecture with Unlabeled Data
- A Comprehensive Survey of Grammar Error Correction
- A Survey of Knowledge Enhanced Pre-trained Models
- Neural Language Generation: Formulation, Methods, and Evaluation
- Context in Neural Machine Translation: A Review of Models and Evaluations
- A neural interlingua for multilingual machine translation
- DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
- Multi-Head Attention with Disagreement Regularization
- Transformer Hawkes Process
- On the adequacy of untuned warmup for adaptive optimization
- A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data Poisoning
- You May Not Need Attention
- Embedding API Dependency Graph for Neural Code Generation
- Neutron: An Implementation of the Transformer Translation Model and its Variants
- BioNetExplorer: Architecture-Space Exploration of Bio-Signal Processing Deep Neural Networks for Wearables
- Trajectory Space Factorization for Deep Video-Based 3D Human Pose Estimation
- An Exploration of Word Embedding Initialization in Deep-Learning Tasks
- STG2Seq: Spatial-temporal Graph to Sequence Model for Multi-step Passenger Demand Forecasting
- Emotion-Aware Transformer Encoder for Empathetic Dialogue Generation
- Collaborative Video Object Segmentation by Foreground-Background Integration
- Distilling Knowledge Learned in BERT for Text Generation
- Highrisk Prediction from Electronic Medical Records via Deep Attention Networks
- Low Resource Neural Machine Translation: A Benchmark for Five African Languages
- TreeBERT: A Tree-Based Pre-Trained Model for Programming Language
- Exploration of Neural Machine Translation in Autoformalization of Mathematics in Mizar
- Review of end-to-end speech synthesis technology based on deep learning
- ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion
- SCAN: Sliding Convolutional Attention Network for Scene Text Recognition
- Multi-granularity Generator for Temporal Action Proposal
- Detecting Hallucinated Content in Conditional Neural Sequence Generation
- Adaptively Sparse Transformers
- Convolutional Self-Attention Networks
- Using deep Residual Networks to search for galaxy-Lyα emitter lens candidates based on spectroscopic-selection
- Towards Robust Neural Machine Translation
- Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings
- Self-Attentive Residual Decoder for Neural Machine Translation
- Learning to Exploit Invariances in Clinical Time-Series Data using Sequence Transformer Networks
- Knowing What, Where and When to Look: Efficient Video Action Modeling with Attention
- Mixed Precision Low-bit Quantization of Neural Network Language Models for Speech Recognition
- Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events
- A Study of Reinforcement Learning for Neural Machine Translation
- Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering
- ConvNets and ImageNet Beyond Accuracy: Understanding Mistakes and Uncovering Biases
- Towards Neural Machine Translation for African Languages
- Non-Autoregressive Neural Machine Translation with Enhanced Decoder Input
- Training Deeper Neural Machine Translation Models with Transparent Attention
- Context-Aware Learning for Neural Machine Translation
- Have convolutions already made recurrence obsolete for unconstrained handwritten text recognition ?
- A Survey on Document-level Neural Machine Translation: Methods and Evaluation
- Compositionality decomposed: how do neural networks generalise?
- Sliced Recurrent Neural Networks
- Cross-Lingual Transfer Learning for Multilingual Task Oriented Dialog
- Attention Aided CSI Wireless Localization
- Time-aware Large Kernel Convolutions
- Neural Machine Translation with Noisy Lexical Constraints
- Pop Music Highlighter: Marking the Emotion Keypoints
- Learning to recognize touch gestures: recurrent vs. convolutional features and dynamic sampling
- A Universal Parent Model for Low-Resource Neural Machine Translation Transfer
- Neural Shuffle-Exchange Networks -- Sequence Processing in O(n log n) Time
- Relative Positional Encoding for Transformers with Linear Complexity
- Breaking the Beam Search Curse: A Study of (Re-)Scoring Methods and Stopping Criteria for Neural Machine Translation
- Unsupervised Neural Machine Translation with Weight Sharing
- Tensor2Tensor for Neural Machine Translation
- Unsupervised Multi-modal Neural Machine Translation
- Deconvolution-Based Global Decoding for Neural Machine Translation
- What they do when in doubt: a study of inductive biases in seq2seq learners
- Information Aggregation for Multi-Head Attention with Routing-by-Agreement
- Gaussian Quadrature for Kernel Features
- Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion
- Denoising Neural Machine Translation Training with Trusted Data and Online Data Selection
- PharmMT: A Neural Machine Translation Approach to Simplify Prescription Directions
- Multimodal Image Captioning for Marketing Analysis
- Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing
- Taming Pretrained Transformers for Extreme Multi-label Text Classification
- The Importance of Being Recurrent for Modeling Hierarchical Structure
- Context- and Sequence-Aware Convolutional Recurrent Encoder for Neural Machine Translation
- Visual Text Correction
- Multilingual Neural Machine Translation with Language Clustering
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta Posterior
- Fast, Diverse and Accurate Image Captioning Guided By Part-of-Speech
- Improving Neural Machine Translation with Conditional Sequence Generative Adversarial Nets
- Hybrid Self-Attention Network for Machine Translation
- A Comprehensive Survey of Deep Learning for Image Captioning
- Source-side Prediction for Neural Headline Generation
- Temporal Deformable Convolutional Encoder-Decoder Networks for Video Captioning
- In Conclusion Not Repetition: Comprehensive Abstractive Summarization With Diversified Attention Based On Determinantal Point Processes
- Semi-Autoregressive Neural Machine Translation
- Abstractive Summarization Using Attentive Neural Techniques
- SING: Symbol-to-Instrument Neural Generator
- Neural Classification of Malicious Scripts: A study with JavaScript and VBScript
- Deep Neural Machine Translation with Weakly-Recurrent Units
- Improve Transformer Models with Better Relative Position Embeddings
- Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient Clipping
- Positively Scale-Invariant Flatness of ReLU Neural Networks
- An Encoder-Decoder Framework Translating Natural Language to Database Queries
- Incremental Adaptation of NMT for Professional Post-editors: A User Study
- Assessment Modeling: Fundamental Pre-training Tasks for Interactive Educational Systems
- Hint-Based Training for Non-Autoregressive Machine Translation
- Towards Understanding Neural Machine Translation with Word Importance
- A Baseline Neural Machine Translation System for Indian Languages
- ZerO Initialization: Initializing Neural Networks with only Zeros and Ones
- ENCORE: Ensemble Learning using Convolution Neural Machine Translation for Automatic Program Repair
- PVRED: A Position-Velocity Recurrent Encoder-Decoder for Human Motion Prediction
- Comprehensible Context-driven Text Game Playing
- LAVA NAT: A Non-Autoregressive Translation Model with Look-Around Decoding and Vocabulary Attention
- Transformer Based Reinforcement Learning For Games
- GRN: Gated Relation Network to Enhance Convolutional Neural Network for Named Entity Recognition
- Defending Against Backdoor Attacks in Natural Language Generation
- A Gap-Based Framework for Chinese Word Segmentation via Very Deep Convolutional Networks
- Dual-Awareness Attention for Few-Shot Object Detection
- Kalman Filtering Attention for User Behavior Modeling in CTR Prediction
- Many-to-Many Voice Transformer Network
- Dynamic Past and Future for Neural Machine Translation
- Hierarchical Pooling Structure for Weakly Labeled Sound Event Detection
- Neural Phrase-to-Phrase Machine Translation
- Automated Classification of Sleep Stages and EEG Artifacts in Mice with Deep Learning
- Double Path Networks for Sequence to Sequence Learning
- Bag-of-Words as Target for Neural Machine Translation
- Regularizing Neural Machine Translation by Target-bidirectional Agreement
- A Study of Multilingual Neural Machine Translation
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
- Regularizing activations in neural networks via distribution matching with the Wasserstein metric
- Lightweight Convolutional Representations for On-Device Natural Language Processing
- Improving Domain Adaptation Translation with Domain Invariant and Specific Information
- Gated Hierarchical Attention for Image Captioning
- Hard but Robust, Easy but Sensitive: How Encoder and Decoder Perform in Neural Machine Translation
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Collaborative Video Object Segmentation by Multi-Scale Foreground-Background Integration
- Neural Chinese Word Segmentation as Sequence to Sequence Translation
- QuickEdit: Editing Text & Translations by Crossing Words Out
- Filtering and Mining Parallel Data in a Joint Multilingual Space
- Convolutional Neural Network for Trajectory Prediction
- LAMBERT: Layout-Aware (Language) Modeling for information extraction
- Constrained Decoding for Neural NLG from Compositional Representations in Task-Oriented Dialogue
- Recurrent Graph Syntax Encoder for Neural Machine Translation
- Multiscale Collaborative Deep Models for Neural Machine Translation
- Improving Conditional Sequence Generative Adversarial Networks by Stepwise Evaluation
- Gated Channel Transformation for Visual Recognition
- Machine Translation : From Statistical to modern Deep-learning practices
- Multi-Domain Adaptation in Neural Machine Translation Through Multidimensional Tagging
- Better Sign Language Translation with STMC-Transformer
- Learning Probabilistic Coordinate Fields for Robust Correspondences
- Neural Machine Translation with Adequacy-Oriented Learning
- Integrative Analysis of Patient Health Records and Neuroimages via Memory-based Graph Convolutional Network
- Dissecting Contextual Word Embeddings: Architecture and Representation
- A General-Purpose Tagger with Convolutional Neural Networks
- Alleviating the Inequality of Attention Heads for Neural Machine Translation
- Future Data Helps Training: Modeling Future Contexts for Session-based Recommendation
- Zero-Resource Neural Machine Translation with Multi-Agent Communication Game
- Beyond BLEU: Training Neural Machine Translation with Semantic Similarity
- Incorporating Word and Subword Units in Unsupervised Machine Translation Using Language Model Rescoring
- PIC: Permutation Invariant Convolution for Recognizing Long-range Activities
- Optimizing Prediction Serving on Low-Latency Serverless Dataflow
- SocialInteractionGAN: Multi-person Interaction Sequence Generation
- Tag-less Back-Translation
- Reflective Decoding Network for Image Captioning
- Greedy Search with Probabilistic N-gram Matching for Neural Machine Translation
- Deep Extreme Multi-label Learning
- Controlled CNN-based Sequence Labeling for Aspect Extraction
- Refining Source Representations with Relation Networks for Neural Machine Translation
- SelfSeg: A Self-supervised Sub-word Segmentation Method for Neural Machine Translation
- Knowledge-Grounded Dialogue Generation with Pre-trained Language Models
- Capacity Control of ReLU Neural Networks by Basis-path Norm
- Would You Ask it that Way? Measuring and Improving Question Naturalness for Knowledge Graph Question Answering
- Optimal Completion Distillation for Sequence Learning
- Robust Neural Malware Detection Models for Emulation Sequence Learning
- Fill in the Blanks: Imputing Missing Sentences for Larger-Context Neural Machine Translation
- Denoising based Sequence-to-Sequence Pre-training for Text Generation
- Tensor Networks for Probabilistic Sequence Modeling
- A Holistic Representation Guided Attention Network for Scene Text Recognition
- A Stable and Effective Learning Strategy for Trainable Greedy Decoding
- Reference Language based Unsupervised Neural Machine Translation
- Cell-aware Stacked LSTMs for Modeling Sentences
- Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention
- Pairwise Interactive Graph Attention Network for Context-Aware Recommendation
- Tensorized Self-Attention: Efficiently Modeling Pairwise and Global Dependencies Together
- Improving Attention Mechanism with Query-Value Interaction
- Contextual Lensing of Universal Sentence Representations
- A Neural Virtual Anchor Synthesizer based on Seq2Seq and GAN Models
- A Simple Convolutional Generative Network for Next Item Recommendation
- Time Distributed Deep Learning Models for Purely Exogenous Forecasting: Application to Water Table Depth Prediction using Weather Image Time Series
- CMFN: Cross-Modal Fusion Network for Irregular Scene Text Recognition
- On-the-Fly Syntax Highlighting using Neural Networks
- Dense Information Flow for Neural Machine Translation
- Multiresolution Transformer Networks: Recurrence is Not Essential for Modeling Hierarchical Structure
- GraphTCN: Spatio-Temporal Interaction Modeling for Human Trajectory Prediction
- Applying SVGD to Bayesian Neural Networks for Cyclical Time-Series Prediction and Inference
- Universal Approximation of Input-Output Maps by Temporal Convolutional Nets
- A Brief Survey of Multilingual Neural Machine Translation
- Neural Architecture Refinement: A Practical Way for Avoiding Overfitting in NAS
- Learning Context-Sensitive Convolutional Filters for Text Processing
- With Little Power Comes Great Responsibility
- Table-to-Text Generation with Effective Hierarchical Encoder on Three Dimensions (Row, Column and Time)
- Learning to Discriminate Noises for Incorporating External Information in Neural Machine Translation
- Refining Source Representations with Relation Networks for Neural Machine Translation
- Joint Detection of Malicious Domains and Infected Clients
- Deep Probabilistic Time Series Forecasting using Augmented Recurrent Input for Dynamic Systems
- Curb Your Carbon Emissions: Benchmarking Carbon Emissions in Machine Translation
- Robust Neural Machine Translation with Joint Textual and Phonetic Embedding
- Fast Prototyping a Dialogue Comprehension System for Nurse-Patient Conversations on Symptom Monitoring
- Transformer-based Arabic Dialect Identification
- Correct Me If You Can: Learning from Error Corrections and Markings
- End-to-End Non-Autoregressive Neural Machine Translation with Connectionist Temporal Classification
- On the Privacy Risks of Deploying Recurrent Neural Networks in Machine Learning Models
- Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention
- Efficient Bidirectional Neural Machine Translation
- Infusing Sequential Information into Conditional Masked Translation Model with Self-Review Mechanism
- Convolutional Neural Networks with Recurrent Neural Filters
- RNNs Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?
- Multi-Task Gaussian Processes and Dilated Convolutional Networks for Reconstruction of Reproductive Hormonal Dynamics
- Knowledge Graph Question Answering via SPARQL Silhouette Generation
- Bayesian Transformer Language Models for Speech Recognition
- Semantic-Unit-Based Dilated Convolution for Multi-Label Text Classification
- A Stochastic Decoder for Neural Machine Translation
- Attentive Convolution: Equipping CNNs with RNN-style Attention Mechanisms
- Title-Guided Encoding for Keyphrase Generation
- Character-Aware Decoder for Translation into Morphologically Rich Languages
- TYolov5: A Temporal Yolov5 Detector Based on Quasi-Recurrent Neural Networks for Real-Time Handgun Detection in Video
- Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning
- Token-level Adaptive Training for Neural Machine Translation
- What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional Encoding
- Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language Generation
- Stock2Vec: A Hybrid Deep Learning Framework for Stock Market Prediction with Representation Learning and Temporal Convolutional Network
- DTMT: A Novel Deep Transition Architecture for Neural Machine Translation
- The TALP-UPC System for the WMT Similar Language Task: Statistical vs Neural Machine Translation
- Dynamic Sentence Sampling for Efficient Training of Neural Machine Translation
- COSEA: Convolutional Code Search with Layer-wise Attention
- Improving Bidirectional Decoding with Dynamic Target Semantics in Neural Machine Translation
- Classifying Long Clinical Documents with Pre-trained Transformers
- Sparse Lifting of Dense Vectors: Unifying Word and Sentence Representations
- Cross-lingual Supervision Improves Unsupervised Neural Machine Translation
- Task-Level Curriculum Learning for Non-Autoregressive Neural Machine Translation
- Attention Focusing for Neural Machine Translation by Bridging Source and Target Embeddings
- Relation Extraction using Explicit Context Conditioning
- Goal-driven Command Recommendations for Analysts
- Neural Machine Translation: A Review and Survey
- Hyperbolic Neural Networks++
- Sequence-Level Knowledge Distillation for Model Compression of Attention-based Sequence-to-Sequence Speech Recognition
- Low Rank Factorization for Compact Multi-Head Self-Attention
- Adversarial Subword Regularization for Robust Neural Machine Translation
- Structure-Invariant Testing for Machine Translation
- Learning Contextualized Sentence Representations for Document-Level Neural Machine Translation
- MUFold-BetaTurn: A Deep Dense Inception Network for Protein Beta-Turn Prediction
- Improving Irregularly Sampled Time Series Learning with Dense Descriptors of Time
- Controllable Dual Skew Divergence Loss for Neural Machine Translation
- Initialization and Regularization of Factorized Neural Layers
- Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers
- Learning Neural Representation of Camera Pose with Matrix Representation of Pose Shift via View Synthesis
- Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks
- Multi-Reference Training with Pseudo-References for Neural Translation and Text Generation
- Once a MAN: Towards Multi-Target Attack via Learning Multi-Target Adversarial Network Once
- Analyzing and Interpreting Convolutional Neural Networks in NLP
- Enhanced Neural Machine Translation by Learning from Draft
- ATCN: Resource-Efficient Processing of Time Series on Edge
- Meta-Embeddings Based On Self-Attention
- SoPa: Bridging CNNs, RNNs, and Weighted Finite-State Machines
- Recommending Multiple Positive Citations for Manuscript via Content-Dependent Modeling and Multi-Positive Triplet
- Lexicon-constrained Copying Network for Chinese Abstractive Summarization
- Shallow-to-Deep Training for Neural Machine Translation
- Normalization Before Shaking Toward Learning Symmetrically Distributed Representation Without Margin in Speech Emotion Recognition
- Listen, Attend, Spell and Adapt: Speaker Adapted Sequence-to-Sequence ASR
- Neural Networks for Predicting Human Interactions in Repeated Games
- STS Classification with Dual-stream CNN
- Iterative Batch Back-Translation for Neural Machine Translation: A Conceptual Model
- Source Dependency-Aware Transformer with Supervised Self-Attention
- Domain Decluttering: Simplifying Images to Mitigate Synthetic-Real Domain Shift and Improve Depth Estimation
- Efficiently Reusing Old Models Across Languages via Transfer Learning
- Converse Attention Knowledge Transfer for Low-Resource Named Entity Recognition
- FPETS : Fully Parallel End-to-End Text-to-Speech System
- A Hierarchical Neural Network for Sequence-to-Sequences Learning
- A Novel Integrated Framework for Learning both Text Detection and Recognition
- Mixture of Expert/Imitator Networks: Scalable Semi-supervised Learning Framework
- TE-ESN: Time Encoding Echo State Network for Prediction Based on Irregularly Sampled Time Series Data
- Language-Independent Representor for Neural Machine Translation
- Machine Translation between Vietnamese and English: an Empirical Study
- Image Captioning as Neural Machine Translation Task in SOCKEYE
- A Semi-Supervised Approach for Abnormal Event Prediction on Large Operational Network Time-Series Data
- Gradients are Not All You Need
- Analyzing Architectures for Neural Machine Translation Using Low Computational Resources
- Rational Recurrences
- A Daily Tourism Demand Prediction Framework Based on Multi-head Attention CNN: The Case of The Foreign Entrant in South Korea
- Learning the Dynamics of Sparsely Observed Interacting Systems
- An Investigation of Warning Erroneous Chat Translations in Cross-lingual Communication
- GWLAN: General Word-Level AutocompletioN for Computer-Aided Translation
- Characters Detection on Namecard with faster RCNN
- Multiplicative Position-aware Transformer Models for Language Understanding
- Transformer-based Automatic Post-Editing with a Context-Aware Encoding Approach for Multi-Source Inputs
- Detecting and Understanding Generalization Barriers for Neural Machine Translation
- Investigating Label Bias in Beam Search for Open-ended Text Generation
- Graph Neural Networks for Node-Level Predictions
- Graph-to-Sequence Neural Machine Translation
- An Interactive Machine Translation Framework for Modernizing Historical Documents
- Shaking Acoustic Spectral Sub-bands Can Better Regularize Learning in Affective Computing
- Generating Diverse Translation by Manipulating Multi-Head Attention
- Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video Data
- A Domain Generalization Perspective on Listwise Context Modeling
- Review-Driven Answer Generation for Product-Related Questions in E-Commerce
- Sequence Generation: From Both Sides to the Middle
- Open-Ended Long-Form Video Question Answering via Hierarchical Convolutional Self-Attention Networks
- Lipschitz Constrained Parameter Initialization for Deep Transformers
- Context-Gated Convolution
- Speeding Up Neural Machine Translation Decoding by Cube Pruning
- Semi-Supervised Few-Shot Learning for Dual Question-Answer Extraction
- Transformer-based Methods for Recognizing Ultra Fine-grained Entities (RUFES)
- Single-Queue Decoding for Neural Machine Translation
- Representing Unordered Data Using Complex-Weighted Multiset Automata
- Is High Variance Unavoidable in RL? A Case Study in Continuous Control
- ARCH: Efficient Adversarial Regularized Training with Caching
- Sequence-to-Sequence Learning with Latent Neural Grammars
- One Model to Learn Both: Zero Pronoun Prediction and Translation
- META: Memory-efficient taxonomic classification and abundance estimation for metagenomics with deep learning
- Neural Machine Translation for Multilingual Grapheme-to-Phoneme Conversion
- Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms
- REAM: An Enhancement Approach to Reference-based Evaluation Metrics for Open-domain Dialog Generation
- Dispatcher: A Message-Passing Approach To Language Modelling
- Adversarial Regularization as Stackelberg Game: An Unrolled Optimization Approach
- Otem&Utem: Over- and Under-Translation Evaluation Metric for NMT
- Automatic Post-Editing for Vietnamese
- GPNAS: A Neural Network Architecture Search Framework Based on Graphical Predictor
- Distributionally Robust Language Modeling
- To Understand Representation of Layer-aware Sequence Encoders as Multi-order-graph
- An Augmented Transformer Architecture for Natural Language Generation Tasks
- Self-Attentional Models Application in Task-Oriented Dialogue Generation Systems
- Token Drop mechanism for Neural Machine Translation
- Cross Copy Network for Dialogue Generation
- Summarizing Videos with Attention
- A Large-Scale Multi-Length Headline Corpus for Analyzing Length-Constrained Headline Generation Model Evaluation
- Learning Hard Retrieval Decoder Attention for Transformers
- Sequential Recommender via Time-aware Attentive Memory Network
- CASA-NLU: Context-Aware Self-Attentive Natural Language Understanding for Task-Oriented Chatbots
- An In-depth Walkthrough on Evolution of Neural Machine Translation
- SongNet: Rigid Formats Controlled Text Generation
- Testing Machine Translation via Referential Transparency
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting
- Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
- Densely Connected Graph Convolutional Networks for Graph-to-Sequence Learning
- All Word Embeddings from One Embedding
- Revisit Systematic Generalization via Meaningful Learning
- GRET: Global Representation Enhanced Transformer
- Regularized Context Gates on Transformer for Machine Translation
- Sentence-wise Smooth Regularization for Sequence to Sequence Learning
- UdS Submission for the WMT 19 Automatic Post-Editing Task
- Self-attention based end-to-end Hindi-English Neural Machine Translation
- Reference Network for Neural Machine Translation
- Latent Part-of-Speech Sequences for Neural Machine Translation
- Improving Multi-Head Attention with Capsule Networks
- Recognizing Arrow Of Time In The Short Stories
- MvSR-NAT: Multi-view Subset Regularization for Non-Autoregressive Machine Translation
- A Deep-Bayesian Framework for Adaptive Speech Duration Modification
- Sentence Structure and Word Relationship Modeling for Emphasis Selection
- DyDiff-VAE: A Dynamic Variational Framework for Information Diffusion Prediction
- A Discriminative Neural Model for Cross-Lingual Word Alignment
- Efficient Purely Convolutional Text Encoding
- Reducing the impact of out of vocabulary words in the translation of natural language questions into SPARQL queries
- Using holistic event information in the trigger
- Exploration into Translation-Equivariant Image Quantization
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Beyond Weight Tying: Learning Joint Input-Output Embeddings for Neural Machine Translation
- Similarity Classification of Public Transit Stations
- Learning a binary search with a recurrent neural network. A novel approach to ordinal regression analysis
- Multi-Glimpse Network: A Robust and Efficient Classification Architecture based on Recurrent Downsampled Attention
- A More Efficient Chinese Named Entity Recognition base on BERT and Syntactic Analysis
- Logographic Subword Model for Neural Machine Translation