68 citations · 97 across the 21 of their papers we have counts for
10 papers · 1 filter
A Proof of Learning Rate Transfer under P
Soufiane Hayou
We provide the first proof of learning rate transfer with width in a linear multi-layer perceptron (MLP) parametrized with P, a neural network parameterization designed to ``max…
Decoupling Dynamical Richness from Representation Learning: Towards Practical Measurement
Yoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee +6
Dynamic feature transformation (the rich regime) does not always align with predictive performance (better representation), yet accuracy is often used as a proxy for richness, limi…
Commutative Width and Depth Scaling in Deep Neural Networks
Soufiane Hayou
This paper is the second in the series Commutative Scaling of Width and Depth (WD) about commutativity of infinite width and depth limits in deep neural networks. Our aim is to und…
On the Connection Between Riemann Hypothesis and a Special Class of Neural Networks
Soufiane Hayou
The Riemann hypothesis (RH) is a long-standing open problem in mathematics. It conjectures that non-trivial zeros of the zeta function all have real part equal to 1/2. The extent o…
Data pruning and neural scaling laws: fundamental limitations of score-based algorithms
Fadhel Ayed, Soufiane Hayou
Data pruning algorithms are commonly used to reduce the memory and computational cost of the optimization process. Recent empirical results reveal that random data pruning remains…
Width and Depth Limits Commute in Residual Networks
Soufiane Hayou, Greg Yang
We show that taking the width and depth to infinity in a deep neural network with skip connections, when branches are scaled by (the only nontrivial scaling), resu…