8 citations · 9 across the 4 of their papers we have counts for
5 papers
Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules
Zhuocheng Gong, Ang Lv, Jian Guan +6
Is it always necessary to compute tokens from shallow to deep layers in Transformers? The continued success of vanilla Transformers and their variants suggests an undoubted "yes".…
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
Bohan Wang, Huishuai Zhang, Qi Meng +3
This paper aims to clearly distinguish between Stochastic Gradient Descent with Momentum (SGDM) and Adam in terms of their convergence rates. We demonstrate that Adam achieves a fa…
Closing the Gap Between the Upper Bound and the Lower Bound of Adam's Iteration Complexity
Bohan Wang, Jingwen Fu, Huishuai Zhang +2
Recently, Arjevani et al. [1] established a lower bound of iteration complexity for the first-order optimization under an -smooth condition and a bounded noise variance assumpti…
Normalized/Clipped SGD with Perturbation for Differentially Private Non-Convex Optimization
Xiaodong Yang, Huishuai Zhang, Wei Chen +1
By ensuring differential privacy in the learning algorithms, one can rigorously mitigate the risk of large models memorizing sensitive training data. In this paper, we study two al…
The Capacity Region of the Source-Type Model for Secret Key and Private Key Generation
Huishuai Zhang, Lifeng Lai, Yingbin Liang +1
The problem of simultaneously generating a secret key (SK) and private key (PK) pair among three terminals via public discussion is investigated, in which each terminal observes a…