24 citations · 28 across the 10 of their papers we have counts for
1 paper · 2 filters
Mostafa Elhoushi, Alex Pretko, Nolan Dey +6
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformer…