597 citations · 1.4k across the 28 of their papers we have counts for
6 papers · 2 filters
Improved Analysis of Clipping Algorithms for Non-convex Optimization
Bohang Zhang, Jikai Jin, Cong Fang +1
Gradient clipping is commonly used in training deep neural networks partly due to its practicability in relieving the exploding gradient problem. Recently, \citet{zhang2019gradient…
Sanity-Checking Pruning Methods: Random Tickets can Win the Jackpot
Jingtong Su, Yihang Chen, Tianle Cai +4
Network pruning is a method for reducing test-time computational resource requirements with minimal performance degradation. Conventional wisdom of pruning algorithms suggests that…
Transferred Discrepancy: Quantifying the Difference Between Representations
Yunzhen Feng, Runtian Zhai, Di He +2
Understanding what information neural networks capture is an essential problem in deep learning, and studying whether different models capture similar features is an initial step t…
(Locally) Differentially Private Combinatorial Semi-Bandits
Xiaoyu Chen, Kai Zheng, Zixin Zhou +3
In this paper, we study Combinatorial Semi-Bandits (CSB) that is an extension of classic Multi-Armed Bandits (MAB) under Differential Privacy (DP) and stronger Local Differential P…
On Layer Normalization in the Transformer Architecture
Ruibin Xiong, Yunchang Yang, Di He +7
The Transformer is widely used in natural language processing tasks. To train a Transformer however, one usually needs a carefully designed learning rate warm-up stage, which is sh…
Combinatorial Semi-Bandit in the Non-Stationary Environment
Wei Chen, Liwei Wang, Haoyu Zhao +1
In this paper, we investigate the non-stationary combinatorial semi-bandit problem, both in the switching case and in the dynamic case. In the general case where (a) the reward fun…