11 citations · 31 across the 8 of their papers we have counts for
8 papers · 1 filter
OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling
Yuxuan Lou, Yang You
Muon fixes the \emph{direction} of every matrix-valued update at the polar factor of its momentum, while each layer's step \emph{magnitude} is addressed only by a static shape corr…
To Repeat or Not To Repeat: Insights from Scaling LLM under Token-Crisis
Fuzhao Xue, Yao Fu, Wangchunshu Zhou +2
Recent research has highlighted the importance of dataset size in scaling language models. However, large language models (LLMs) are notoriously token-hungry during pre-training, a…
Towards Efficient and Scalable Sharpness-Aware Minimization
Yong Liu, Siqi Mai, Xiangning Chen +2
Recently, Sharpness-Aware Minimization (SAM), which connects the geometry of the loss landscape and generalization, has demonstrated significant performance boosts on training larg…
Large-Scale Deep Learning Optimizations: A Comprehensive Survey
Xiaoxin He, Fuzhao Xue, Xiaozhe Ren +1
Deep learning have achieved promising results on a wide spectrum of AI applications. Larger datasets and models consistently yield better performance. However, we generally spend l…
Go Wider Instead of Deeper
Fuzhao Xue, Ziji Shi, Futao Wei +3
More transformer blocks with residual connections have recently achieved impressive results on various tasks. To achieve better performance with fewer trainable parameters, recent…
Training EfficientNets at Supercomputer Scale: 83% ImageNet Top-1 Accuracy in One Hour
Arissa Wongpanich, Hieu Pham, James Demmel +4
EfficientNets are a family of state-of-the-art image classification models based on efficiently scaled convolutional neural networks. Currently, EfficientNets can take on the order…