28 citations · 28 across the 1 of their papers we have counts for
2 papers
cs.CL2021★ 28 cited
NormFormer: Improved Transformer Pretraining with Extra Normalization
Sam Shleifer, Jason Weston, Myle Ott
During pretraining, the Pre-LayerNorm transformer suffers from a gradient magnitude mismatch: gradients at early layers are much larger than at later layers. These issues can be al…
cs.CL2019
HuggingFace's Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh +19
Recent progress in natural language processing has been driven by advances in both model architecture and model pretraining. Transformer architectures have facilitated building hig…