1 paper
Inyoung Paik, Jaesik Choi
Deep neural networks, which employ batch normalization and ReLU-like activation functions, suffer from instability in the early stages of training due to the high gradient induced…