1 paper · 1 filter
Peng Xu, Dhruv Kumar, Wei Yang +6
It is a common belief that training deep transformers from scratch requires large datasets. Consequently, for small datasets, people usually use shallow and simple additional layer…