From the 1 of 5 linked papers with an AI index.
5 papers
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
Tiberiu Musat, Tiago Pimentel, Nicolas Zucchet +1
The paper develops a theoretical framework showing that transformer models learn inductive reasoning tasks by evolving on a low-dimensional invariant manifold, enabling tractable a…
The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold
Tiberiu Musat
Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data. Previ…
Neural Weight Norm = Kolmogorov Complexity
Tiberiu Musat
Why does weight decay work? We prove that, in any fixed-precision regime, the smallest weight norm of a looped neural network outputting a binary string equals the Kolmogorov compl…
On the Emergence of Induction Heads for In-Context Learning
Tiberiu Musat, Tiago Pimentel, Lorenzo Noci +3
Transformers have become the dominant architecture for natural language processing. Part of their success is owed to a remarkable capability known as in-context learning (ICL): the…
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
Tiberiu Musat
In this paper, I introduce the retrieval problem, a simple yet common reasoning task that can be solved only by transformers with a minimum number of layers, which grows logarithmi…