17 papers
Disentangling the Expressivity of RoPE
Selim Jerad, Anej Svete, Jiaoda Li +1
Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, wherea…
Generating in the Limit with Infinitely Many Hallucinations
Irene Strauss, Alexandra Butoi, Ryan Cotterell
The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner task…
Causally Evaluating the Learnability of Formal Language Tasks
Vésteinn Snæbjarnarson, Anej Svete, Josef Valvoda +3
Language models, as multi-task learners, acquire a wide range of abilities during training. A fundamental question is how much task-specific data is needed to learn a given task. A…
Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions
Blanka Köver, Alexandra Butoi, Anej Svete +2
Transformers consistently fail to learn certain simple functions that are provably expressible with specific parameter settings. This gap between learnability and expressivity is p…
Context-Free Recognition with Transformers
Selim Jerad, Anej Svete, Sophie Hao +2
Transformers excel empirically on tasks that process well-formed inputs according to some grammar, such as natural language and code. However, it remains unclear how they can proce…
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
Anej Svete, William Merrill, Ryan Cotterell +1
Recent work describes what transformers can and cannot compute through connections to boolean circuits, but existing results lack exact characterizations and are sensitive to model…