4 papers
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization
Felipe Urrutia, Juan José AlegrÃa, Cinthia Sanchez Macias +3
Transformer-based language models are widespread in today's society. As such, understanding the mechanisms by which they solve structured tasks and predicting how they may behave i…
Message Passing on the Edge: Towards Scalable and Expressive GNNs
Pablo Barceló, Fabian Jogl, Alexander Kozachinskiy +3
Graph neural networks (GNNs) are widely used in graph learning and most architectures propagate information by passing messages between vertices. In this work, we shift our attenti…
Decoupling Positional and Symbolic Attention Behavior in Transformers
Felipe Urrutia, Jorge Salas, Alexander Kozachinskiy +3
An important aspect subtending language understanding and production is the ability to independently encode positional and symbolic information of the words within a sentence. In T…
Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models
Hector Pasten, Felipe Urrutia, Hector Jimenez +3
Understanding how Transformers work and how they process information is key to the theoretical and empirical advancement of these machines. In this work, we demonstrate the existen…