11 papers
Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence
Itay Lavie, Kirsten Fischer, Andrey Lekov +3
Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present…
Improving CFT Operators Using Machine Learning
Lior Oppenheim, Snir Gazit, Zohar Ringel
Finite-size effects limit the accuracy with which conformal data can be extracted from lattice simulations of critical systems. While action improvement suppresses some corrections…
Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers
Orit Davidovich, Zohar Ringel
We formally define algorithmic capture of combinatorial tasks as the ability of a transformer to extrapolate to arbitrary task sizes with controllable error and logarithmic sample…
Mitigating the Curse of Detail: Scaling Arguments for Feature Learning and Sample Complexity
Noa Rubin, Orit Davidovich, Zohar Ringel
Two pressing topics in the theory of deep learning are the interpretation of feature learning (FL) mechanisms and the determination of implicit bias of networks in the rich regime.…
Lecture notes: From Gaussian processes to feature learning
Moritz Helias, Javed Lindner, Lars Schutzeichel +1
These lecture notes develop the theory of learning in deep and recurrent neuronal networks from the point of view of Bayesian inference. The aim is to enable the reader to understa…
Renormalization group for deep neural networks: Universality of learning and scaling laws
Gorka Peraza Coppola, Moritz Helias, Zohar Ringel
Self-similarity, where observables at different length scales exhibit similar behavior, is ubiquitous in natural systems. Such systems are typically characterized by power-law corr…