3 citations · 3 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
Gavia Gray, Aman Tiwari, Shane Bergsma +1
Per-example gradient norms are a vital ingredient for estimating gradient noise scale (GNS) with minimal variance. Observing the tensor contractions required to compute them, we pr…
cs.LG2019
BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget
Jack Turner, Elliot J. Crowley, Michael O'Boyle +2
The desire to map neural networks to varying-capacity devices has led to the development of a wealth of compression techniques, many of which involve replacing standard convolution…