10 citations · 17 across the 3 of their papers we have counts for
3 papers · 1 filter
Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization
Stanislaw Jastrzebski, Devansh Arpit, Oliver Astrand +6
The early phase of training a deep neural network has a dramatic effect on the local curvature of the loss function. For instance, using a small learning rate does not guarantee st…
Advantages of biologically-inspired adaptive neural activation in RNNs during learning
Victor Geadah, Giancarlo Kerg, Stefan Horoi +2
Dynamic adaptation in single-neuron response plays a fundamental role in neural coding in biological neural networks. Yet, most neural activation functions used in artificial netwo…
Untangling tradeoffs between recurrence and self-attention in neural networks
Giancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal +3
Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with li…