4 papers
Conda: Column-Normalized Adam for Training Large Language Models Faster
Junjie Wang, Pan Zhou, Yiming Dong +6
Large language models (LLMs) have demonstrated impressive generalization and emergent capabilities, yet their pre-training remains computationally expensive and sensitive to optimi…
On the Convergence Rate of AdamW Measured by Norm
Huan Li, Yiming Dong, Zhouchen Lin
As the default optimizer for training large language models, AdamW has achieved remarkable success in deep learning. However, its convergence behavior is not theoretically well-und…
Revisiting EXTRA for Smooth Distributed Optimization
Huan Li, Zhouchen Lin
EXTRA is a popular method for dencentralized distributed optimization and has broad applications. This paper revisits EXTRA. First, we give a sharp complexity analysis for EXTRA wi…
Decentralized Accelerated Gradient Methods With Increasing Penalty Parameters
Huan Li, Cong Fang, Wotao Yin +1
In this paper, we study the communication and (sub)gradient computation costs in distributed optimization and give a sharp complexity analysis for the proposed distributed accelera…