4 papers
Sequential Group Composition: A Window into the Mechanics of Deep Learning
Giovanni Luca Marchetti, Daniel Kunin, Adele Myers +2
How do neural networks trained over sequences acquire the ability to perform structured operations, such as arithmetic, geometric, and algorithmic computation? To gain insight into…
Symmetry Breaking in Transformers for Efficient and Interpretable Training
Eva Silverstein, Daniel Kunin, Vasudev Shyam
The attention mechanism in its standard implementation contains extraneous rotational degrees of freedom that are carried through computation but do not affect model activations or…
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
Daniel Kunin, Giovanni Luca Marchetti, Feng Chen +5
What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dy…
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
Clémentine C. J. Dominé, Nicolas Anguita, Alexandra M. Proca +4
Biological and artificial neural networks develop internal representations that enable them to perform complex tasks. In artificial networks, the effectiveness of these models reli…