Exploring Fine-tuning Techniques for Pre-trained Cross-lingual Models via Continual Learning
arXiv:2004.14218
Abstract
Recently, fine-tuning pre-trained language models (e.g., multilingual BERT) to downstream cross-lingual tasks has shown promising results. However, the fine-tuning process inevitably changes the parameters of the pre-trained model and weakens its cross-lingual ability, which leads to sub-optimal performance. To alleviate this problem, we leverage continual learning to preserve the original cross-lingual ability of the pre-trained model when we fine-tune it to downstream tasks. The experimental result shows that our fine-tuning methods can better preserve the cross-lingual ability of the pre-trained model in a sentence retrieval task. Our methods also achieve better performance than other fine-tuning baselines on the zero-shot cross-lingual part-of-speech tagging and named entity recognition tasks.
References in corpus (5)
- Word Translation Without Parallel Data
- Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks
- XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation
- Attention-Informed Mixed-Language Training for Zero-shot Cross-lingual Task-oriented Dialogue Systems
- Coach: A Coarse-to-Fine Approach for Cross-domain Slot Filling