178 citations · 343 across the 10 of their papers we have counts for
4 papers · 1 filter
Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)
Yu Huang, Junyang Lin, Chang Zhou +2
Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network ou…
M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining
Junyang Lin, An Yang, Jinze Bai +9
Recent expeditious developments in deep learning algorithms, distributed training, and even hardware design for large models have enabled training extreme-scale models, say GPT-3 a…
M6-T: Exploring Sparse Expert Models and Beyond
An Yang, Junyang Lin, Rui Men +12
Mixture-of-Experts (MoE) models can achieve promising results with outrageous large amount of parameters but constant computation cost, and thus it has become a trend in model scal…
Understanding and Improving Layer Normalization
Jingjing Xu, Xu Sun, Zhiyuan Zhang +2
Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accu…