1 citations · 1 across the 1 of their papers we have counts for
1 paper
Asher Trockman, J. Zico Kolter
It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-…