353 citations · 360 across the 2 of their papers we have counts for
3 papers
Variational Bi-LSTMs
Samira Shabanian, Devansh Arpit, Adam Trischler +1
Recurrent neural networks like long short-term memory (LSTM) are important architectures for sequential prediction tasks. LSTMs (and RNNs in general) model sequences along the forw…
A Closer Look at Memorization in Deep Networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas +8
We examine the role of memorization in deep learning, drawing connections to capacity, generalization, and adversarial robustness. While deep networks are capable of memorizing noi…
Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep Networks
Devansh Arpit, Yingbo Zhou, Bhargava U. Kota +1
While the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks-- Internal Covariate Shift-- the current solution has cer…