4 papers
The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis
Hoang Pham, The-Anh Ta, Tom Jacobs +2
Sparse neural networks promise efficiency, yet training them effectively remains a fundamental challenge. Despite advances in pruning methods that create sparse architectures, unde…
Pay Attention to Small Weights
Chao Zhou, Tom Jacobs, Advait Gadhikar +1
Finetuning large pretrained neural networks is known to be resource-intensive, both in terms of memory and computational cost. To mitigate this, a common approach is to restrict tr…
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
Tom Jacobs, Chao Zhou, Rebekka Burkholz
Implicit bias plays an important role in explaining how overparameterized models generalize well. Explicit regularization like weight decay is often employed in addition to prevent…
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
Advait Gadhikar, Tom Jacobs, Chao Zhou +1
The performance gap between training sparse neural networks from scratch (PaI) and dense-to-sparse training presents a major roadblock for efficient deep learning. According to the…