Analyzing Architectures for Neural Machine Translation Using Low Computational Resources
arXiv:2111.03813 · doi:10.5121/ijnlc.2021.10502
Abstract
With the recent developments in the field of Natural Language Processing, there has been a rise in the use of different architectures for Neural Machine Translation. Transformer architectures are used to achieve state-of-the-art accuracy, but they are very computationally expensive to train. Everyone cannot have such setups consisting of high-end GPUs and other resources. We train our models on low computational resources and investigate the results. As expected, transformers outperformed other architectures, but there were some surprising results. Transformers consisting of more encoders and decoders took more time to train but had fewer BLEU scores. LSTM performed well in the experiment and took comparatively less time to train than transformers, making it suitable to use in situations having time constraints.
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- SYSTRAN's Pure Neural Machine Translation Systems
- PMIndia -- A Collection of Parallel Corpora of Languages of India
- WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia
- Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages
- Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
- Marathi To English Neural Machine Translation With Near Perfect Corpus And Transformers