1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Tao Yang, Jinghao Deng, Xiaojun Quan +2
Fine-tuning large pre-trained language models on downstream tasks is apt to suffer from overfitting when limited training data is available. While dropout proves to be an effective…