5 papers
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization
Felipe Urrutia, Juan José AlegrÃa, Cinthia Sanchez Macias +3
Transformer-based language models are widespread in today's society. As such, understanding the mechanisms by which they solve structured tasks and predicting how they may behave i…
Decoupling Positional and Symbolic Attention Behavior in Transformers
Felipe Urrutia, Jorge Salas, Alexander Kozachinskiy +3
An important aspect subtending language understanding and production is the ability to independently encode positional and symbolic information of the words within a sentence. In T…
Strassen Attention, Split VC Dimension and Compositionality in Transformers
Alexander Kozachinskiy, Felipe Urrutia, Hector Jimenez +6
We propose the first method to show theoretical limitations for one-layer softmax transformers with arbitrarily many precision bits (even infinite). We establish those limitations…
Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models
Hector Pasten, Felipe Urrutia, Hector Jimenez +3
Understanding how Transformers work and how they process information is key to the theoretical and empirical advancement of these machines. In this work, we demonstrate the existen…
Gradient-based inference of abstract task representations for generalization in neural networks
Ali Hummos, Felipe del RÃo, Brabeeba Mien Wang +3
Humans and many animals show remarkably adaptive behavior and can respond differently to the same input depending on their internal goals. The brain not only represents the interme…