IndicBART: A Pre-trained Model for Indic Natural Language Generation
arXiv:2109.02903 · doi:10.18653/v1/2022.findings-acl.145
Abstract
In this paper, we study pre-trained sequence-to-sequence models for a group of related languages, with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages and English. IndicBART utilizes the orthographic similarity between Indic scripts to improve transfer learning between similar Indic languages. We evaluate IndicBART on two NLG tasks: Neural Machine Translation (NMT) and extreme summarization. Our experiments on NMT and extreme summarization show that a model specific to related languages like IndicBART is competitive with large pre-trained models like mBART50 despite being significantly smaller. It also performs well on very low-resource translation scenarios where languages are not included in pre-training or fine-tuning. Script sharing, multilingual training, and better utilization of limited model capacity contribute to the good performance of the compact IndicBART model.
Published at ACL 2022, 15 pages
References in corpus (8)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- MuRIL: Multilingual Representations for Indian Languages
- Transfer Learning across Low-Resource, Related Languages for Neural Machine Translation
- The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
- PMIndia -- A Collection of Parallel Corpora of Languages of India
- A Multilingual Parallel Corpora Collection Effort for Indian Languages
- YANMTT: Yet Another Neural Machine Translation Toolkit