54 citations · 73 across the 3 of their papers we have counts for
3 papers
cs.LG2022★ 6 cited
TorchScale: Transformers at Scale
Shuming Ma, Hongyu Wang, Shaohan Huang +8
Large Transformers have achieved state-of-the-art performance across many tasks. Most open-source libraries on scaling Transformers focus on improving training or inference with be…
cs.LG2022★ 13 cited
Foundation Transformers
Hongyu Wang, Shuming Ma, Shaohan Huang +12
A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different imp…
cs.CL2022★ 54 cited
DeepNet: Scaling Transformers to 1,000 Layers
Hongyu Wang, Shuming Ma, Li Dong +3
In this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the r…