118 citations · 511 across the 21 of their papers we have counts for
4 papers · 2 filters
Trust Region Based Adversarial Attack on Neural Networks
Zhewei Yao, Amir Gholami, Peng Xu +2
Deep Neural Networks are quite vulnerable to adversarial perturbations. Current state-of-the-art adversarial attack methods typically require very time consuming hyper-parameter tu…
Parameter Re-Initialization through Cyclical Batch Size Schedules
Norman Mu, Zhewei Yao, Amir Gholami +2
Optimal parameter initialization remains a crucial problem for neural network training. A poor weight initialization may take longer to train and/or converge to sub-optimal solutio…
On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
Noah Golmant, Nikita Vemuri, Zhewei Yao +5
Increasing the mini-batch size for stochastic gradient descent offers significant opportunities to reduce wall-clock training time, but there are a variety of theoretical and syste…
Large batch size training of neural networks with adversarial training and second-order information
Zhewei Yao, Amir Gholami, Daiyaan Arfeen +4
The most straightforward method to accelerate Stochastic Gradient Descent (SGD) computation is to distribute the randomly selected batch of inputs over multiple processors. To keep…