12 papers
Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence
Itay Lavie, Kirsten Fischer, Andrey Lekov +3
Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present…
Discrete signaling mediates chaotic regularization in recurrent neural networks
Jan Bauer, Christian Keup, Jonathan Kadmon +1
Cortical circuits operate in a regime of intrinsic chaos, where even tiny changes in input can lead to divergent neural responses. Yet, remarkably, population codes in the brain va…
Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues
Jakob Kramp, Javed Lindner, Moritz Helias
Training large neural networks exposes neural scaling laws for the generalization error, which points to a universal behavior across network architectures of learning in high dimen…
A unified theory of feature learning in RNNs and DNNs
Jan P. Bauer, Kirsten Fischer, Moritz Helias +1
Recurrent and deep neural networks (RNNs/DNNs) are cornerstone architectures in machine learning. Remarkably, RNNs differ from DNNs only by weight sharing, as can be shown through…
Lecture notes: From Gaussian processes to feature learning
Moritz Helias, Javed Lindner, Lars Schutzeichel +1
These lecture notes develop the theory of learning in deep and recurrent neuronal networks from the point of view of Bayesian inference. The aim is to enable the reader to understa…
Renormalization group for deep neural networks: Universality of learning and scaling laws
Gorka Peraza Coppola, Moritz Helias, Zohar Ringel
Self-similarity, where observables at different length scales exhibit similar behavior, is ubiquitous in natural systems. Such systems are typically characterized by power-law corr…