activity
20242026
collaborators

19 papers

cs.LG2026

Disentangling the Expressivity of RoPE

Selim Jerad, Anej Svete, Jiaoda Li +1

Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, wherea…

cs.CL2026

Generating in the Limit with Infinitely Many Hallucinations

Irene Strauss, Alexandra Butoi, Ryan Cotterell

The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner task…

cs.CL2026

Causally Evaluating the Learnability of Formal Language Tasks

Vésteinn Snæbjarnarson, Anej Svete, Josef Valvoda +3

Language models, as multi-task learners, acquire a wide range of abilities during training. A fundamental question is how much task-specific data is needed to learn a given task. A…

cs.LG2026

Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions

Blanka Köver, Alexandra Butoi, Anej Svete +2

Transformers consistently fail to learn certain simple functions that are provably expressible with specific parameter settings. This gap between learnability and expressivity is p…

cs.LG2026

Context-Free Recognition with Transformers

Selim Jerad, Anej Svete, Sophie Hao +2

Transformers excel empirically on tasks that process well-formed inputs according to some grammar, such as natural language and code. However, it remains unclear how they can proce…

cs.LG2026

Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't

Anej Svete, William Merrill, Ryan Cotterell +1

Recent work describes what transformers can and cannot compute through connections to boolean circuits, but existing results lack exact characterizations and are sensitive to model…