2 citations · 2 across the 16 of their papers we have counts for
13 papers · 1 filter
Indexing: the Beginning and the End
Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia
We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of th…
Parity, Sensitivity, and Transformers
Alexander Kozachinskiy, Tomasz Steifer, Przemysław Wałȩga
Understanding what neural architectures can and cannot compute is a central challenge in the theory of AI. One of the fundamental problems in this context is the PARITY task, which…
Decoupling Positional and Symbolic Attention Behavior in Transformers
Felipe Urrutia, Jorge Salas, Alexander Kozachinskiy +3
An important aspect subtending language understanding and production is the ability to independently encode positional and symbolic information of the words within a sentence. In T…
Message Passing on the Edge: Towards Scalable and Expressive GNNs
Pablo Barceló, Fabian Jogl, Alexander Kozachinskiy +3
Graph neural networks (GNNs) are widely used in graph learning and most architectures propagate information by passing messages between vertices. In this work, we shift our attenti…
Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models
Hector Pasten, Felipe Urrutia, Hector Jimenez +3
Understanding how Transformers work and how they process information is key to the theoretical and empirical advancement of these machines. In this work, we demonstrate the existen…
A completely uniform transformer for parity
Alexander Kozachinskiy, Tomasz Steifer
We construct a 3-layer constant-dimension transformer, recognizing the parity language, where neither parameter matrices nor the positional encoding depend on the input length. Thi…