Sockeye: A Toolkit for Neural Machine Translation
arXiv:1712.05690
Abstract
We describe Sockeye (version 1.12), an open-source sequence-to-sequence toolkit for Neural Machine Translation (NMT). Sockeye is a production-ready framework for training and applying models as well as an experimental platform for researchers. Written in Python and built on MXNet, the toolkit offers scalable training and inference for the three most prominent encoder-decoder architectures: attentional recurrent neural networks, self-attentional transformers, and fully convolutional networks. Sockeye also supports a wide range of optimizers, normalization and regularization techniques, and inference improvements from current NMT literature. Users can easily run standard training recipes, explore different model settings, and incorporate new ideas. In this paper, we highlight Sockeye's features and benchmark it against other NMT toolkits on two language arcs from the 2017 Conference on Machine Translation (WMT): English-German and Latvian-English. We report competitive BLEU scores across all three architectures, including an overall best score for Sockeye's transformer implementation. To facilitate further comparison, we release all system outputs and training scripts used in our experiments. The Sockeye toolkit is free software released under the Apache 2.0 license.
References in corpus (1)
Cited by in corpus (62)
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- An Empirical Exploration of Curriculum Learning for Neural Machine Translation
- OpenNMT: Neural Machine Translation Toolkit
- RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition
- Freezing Subnetworks to Analyze Domain Adaptation in Neural Machine Translation
- Multi-Task Neural Models for Translating Between Styles Within and Across Languages
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- Physics-guided Deep Markov Models for Learning Nonlinear Dynamical Systems with Uncertainty
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- Priority-based Parameter Propagation for Distributed DNN Training
- Robust Neural Machine Translation with Doubly Adversarial Inputs
- Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
- Learning to Segment Inputs for NMT Favors Character-Level Processing
- Context in Neural Machine Translation: A Review of Models and Evaluations
- Encoders Help You Disambiguate Word Senses in Neural Machine Translation
- Explaining Sequence-Level Knowledge Distillation as Data-Augmentation for Neural Machine Translation
- Neutron: An Implementation of the Transformer Translation Model and its Variants
- Morphological Word Segmentation on Agglutinative Languages for Neural Machine Translation
- Effective Cross-lingual Transfer of Neural Machine Translation Models without Shared Vocabularies
- Neural Polysynthetic Language Modelling
- Bi-Directional Neural Machine Translation with Synthetic Parallel Data
- Marathi To English Neural Machine Translation With Near Perfect Corpus And Transformers
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Grammatical Error Correction and Style Transfer via Zero-shot Monolingual Translation
- Machine Translation System Selection from Bandit Feedback
- Problems with automating translation of movie/TV show subtitles
- Curriculum Learning for Domain Adaptation in Neural Machine Translation
- Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN Training
- Document-aligned Japanese-English Conversation Parallel Corpus
- Complexity-Weighted Loss and Diverse Reranking for Sentence Simplification
- Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
- An Analysis of Attention Mechanisms: The Case of Word Sense Disambiguation in Neural Machine Translation
- Learning Bilingual Sentence Embeddings via Autoencoding and Computing Similarities with a Multilayer Perceptron
- End-to-End Non-Autoregressive Neural Machine Translation with Connectionist Temporal Classification
- Addressing Zero-Resource Domains Using Document-Level Context in Neural Machine Translation
- A Stochastic Decoder for Neural Machine Translation
- Controlling Text Complexity in Neural Machine Translation
- Extremely low-resource machine translation for closely related languages
- Neural Machine Translation: A Review and Survey
- Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English
- Monolingual and Cross-lingual Zero-shot Style Transfer
- Graph-to-Sequence Learning using Gated Graph Neural Networks
- Differentiable Sampling with Flexible Reference Word Order for Neural Machine Translation
- DeepSubQE: Quality estimation for subtitle translations
- Image Captioning as Neural Machine Translation Task in SOCKEYE
- Lipschitz Constrained Parameter Initialization for Deep Transformers
- Sicilian Translator: A Recipe for Low-Resource NMT
- Bi-Directional Differentiable Input Reconstruction for Low-Resource Neural Machine Translation
- Data Ordering Patterns for Neural Machine Translation: An Empirical Study
- Evaluating Robustness to Input Perturbations for Neural Machine Translation
- Neural Machine Translation for Multilingual Grapheme-to-Phoneme Conversion
- Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text Generation
- Designing the Business Conversation Corpus
- Revisiting Negation in Neural Machine Translation
- Controlling Neural Machine Translation Formality with Synthetic Supervision
- Improving Zero-shot Multilingual Neural Machine Translation for Low-Resource Languages
- A Discriminative Neural Model for Cross-Lingual Word Alignment
- uniblock: Scoring and Filtering Corpus with Unicode Block Information
- EDITOR: an Edit-Based Transformer with Repositioning for Neural Machine Translation with Soft Lexical Constraints
- Improving Unsupervised Word-by-Word Translation with Language Model and Denoising Autoencoder
- Densely Connected Graph Convolutional Networks for Graph-to-Sequence Learning
- Understanding Neural Machine Translation by Simplification: The Case of Encoder-free Models