1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Dezhou Shen
The mainstream BERT/GPT model contains only 10 to 20 layers, and there is little literature to discuss the training of deep BERT/GPT. This paper proposes a simple yet effective met…