activity
20172026
most citedOn the Impact of the Activation Function on Deep Neural Networks Training

68 citations · 97 across the 21 of their papers we have counts for

collaborators
Showing stat.MLShow all

10 papers · 1 filter

stat.ML2025

A Proof of Learning Rate Transfer under P

Soufiane Hayou

We provide the first proof of learning rate transfer with width in a linear multi-layer perceptron (MLP) parametrized with P, a neural network parameterization designed to ``max…

stat.ML2024

Decoupling Dynamical Richness from Representation Learning: Towards Practical Measurement

Yoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee +6

Dynamic feature transformation (the rich regime) does not always align with predictive performance (better representation), yet accuracy is often used as a proxy for richness, limi…

stat.ML2023

Commutative Width and Depth Scaling in Deep Neural Networks

Soufiane Hayou

This paper is the second in the series Commutative Scaling of Width and Depth (WD) about commutativity of infinite width and depth limits in deep neural networks. Our aim is to und…

stat.ML2023

On the Connection Between Riemann Hypothesis and a Special Class of Neural Networks

Soufiane Hayou

The Riemann hypothesis (RH) is a long-standing open problem in mathematics. It conjectures that non-trivial zeros of the zeta function all have real part equal to 1/2. The extent o…

stat.ML2023

Data pruning and neural scaling laws: fundamental limitations of score-based algorithms

Fadhel Ayed, Soufiane Hayou

Data pruning algorithms are commonly used to reduce the memory and computational cost of the optimization process. Recent empirical results reveal that random data pruning remains…

stat.ML2023★ 1 cited

Width and Depth Limits Commute in Residual Networks

Soufiane Hayou, Greg Yang

We show that taking the width and depth to infinity in a deep neural network with skip connections, when branches are scaled by (the only nontrivial scaling), resu…