A Focus on Neural Machine Translation for African Languages
arXiv:1906.05685
Abstract
African languages are numerous, complex and low-resourced. The datasets required for machine translation are difficult to discover, and existing research is hard to reproduce. Minimal attention has been given to machine translation for African languages so there is scant research regarding the problems that arise when using machine translation techniques. To begin addressing these problems, we trained models to translate English to five of the official South African languages (Afrikaans, isiZulu, Northern Sotho, Setswana, Xitsonga), making use of modern neural machine translation techniques. The results obtained show the promise of using neural machine translation techniques for African languages. By providing reproducible publicly-available data, code and results, this research aims to provide a starting point for other researchers in African machine translation to compare to and build upon.
References in corpus (1)
Cited by in corpus (6)
- Benchmarking Neural Machine Translation for Southern African Languages
- NLP for Ghanaian Languages
- Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages
- Lanfrica: A Participatory Approach to Documenting Machine Translation Research on African Languages
- Neural Machine Translation for Extremely Low-Resource African Languages: A Case Study on Bambara
- The Low-Resource Double Bind: An Empirical Study of Pruning for Low-Resource Machine Translation