14 citations · 27 across the 7 of their papers we have counts for
8 papers
Parameter and Data Efficient Continual Pre-training for Robustness to Dialectal Variance in Arabic
Soumajyoti Sarkar, Kaixiang Lin, Sailik Sengupta +3
The use of multilingual language models for tasks in low and high-resource languages has been a success story in deep learning. In recent times, Arabic has been receiving widesprea…
Distiller: A Systematic Study of Model Distillation Methods in Natural Language Processing
Haoyu He, Xingjian Shi, Jonas Mueller +3
We aim to identify how different components in the KD pipeline affect the resulting performance and how much the optimal KD pipeline varies across different datasets/tasks, such as…
Accelerated Large Batch Optimization of BERT Pretraining in 54 minutes
Shuai Zheng, Haibin Lin, Sheng Zha +1
BERT has recently attracted a lot of attention in natural language understanding (NLU) and achieved state-of-the-art results in various NLU tasks. However, its success requires lar…
Unlearn Dataset Bias in Natural Language Inference by Fitting the Residual
He He, Sheng Zha, Haohan Wang
Statistical natural language inference (NLI) models are susceptible to learning dataset bias: superficial cues that happen to associate with the label on a particular dataset, but…
GluonCV and GluonNLP: Deep Learning in Computer Vision and Natural Language Processing
Jian Guo, He He, Tong He +13
We present GluonCV and GluonNLP, the deep learning toolkits for computer vision and natural language processing based on Apache MXNet (incubating). These toolkits provide state-of-…
Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
Haibin Lin, Hang Zhang, Yifei Ma +4
With an increasing demand for training powers for deep learning algorithms and the rapid growth of computation resources in data centers, it is desirable to dynamically schedule di…