1 paper
Vignesh Kothapalli, Tianyu Pang, Shenyang Deng +2
Training strategies for modern deep neural networks (NNs) tend to induce a heavy-tailed (HT) empirical spectral density (ESD) in the layer weights. While previous efforts have show…