activity
20242026
collaborators

5 papers

cs.LG2026

Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization

Felipe Urrutia, Juan José Alegría, Cinthia Sanchez Macias +3

Transformer-based language models are widespread in today's society. As such, understanding the mechanisms by which they solve structured tasks and predicting how they may behave i…

cs.LG2025

Decoupling Positional and Symbolic Attention Behavior in Transformers

Felipe Urrutia, Jorge Salas, Alexander Kozachinskiy +3

An important aspect subtending language understanding and production is the ability to independently encode positional and symbolic information of the words within a sentence. In T…

cs.LG2025

Strassen Attention, Split VC Dimension and Compositionality in Transformers

Alexander Kozachinskiy, Felipe Urrutia, Hector Jimenez +6

We propose the first method to show theoretical limitations for one-layer softmax transformers with arbitrarily many precision bits (even infinite). We establish those limitations…

cs.LG2025

Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models

Hector Pasten, Felipe Urrutia, Hector Jimenez +3

Understanding how Transformers work and how they process information is key to the theoretical and empirical advancement of these machines. In this work, we demonstrate the existen…

cs.LG2024

Gradient-based inference of abstract task representations for generalization in neural networks

Ali Hummos, Felipe del Río, Brabeeba Mien Wang +3

Humans and many animals show remarkably adaptive behavior and can respond differently to the same input depending on their internal goals. The brain not only represents the interme…