From the 1 of 5 linked papers with an AI index.
5 papers
A Compositional Theory of Causally Masked Transformers
Franz Nowak, Ryan Cotterell, Reda Boumasmoud
The paper develops an algebraic framework to characterize what decision problems finite‑precision, causally masked transformers can solve, linking attention mechanisms to memory re…
Characterizing the Expressivity of Local Attention in Transformers
Jiaoda Li, Ryan Cotterell
The transformer is the most popular neural architecture for language modeling. The cornerstone of the transformer is its global attention mechanism, which lets the model aggregate…
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
Brian DuSell, Ryan Cotterell
When children learn language, they make syntactic generalizations based on hierarchical rules. A recent line of work has inquired as to whether common neural network architectures…
An Algebraic View of the Expressivity of Recurrent Language Models
Franz Nowak, Ryan Cotterell, Reda Boumasmoud
What formal languages can a recurrent neural language model recognize? Formal results in the literature conflict: some authors report Turing-completeness, while others show equival…
Transformers are Inherently Succinct
Pascal BergsträÃer, Ryan Cotterell, Anthony W. Lin
We study succinctness as a measure of the expressive power of transformers. Succinctness -- how compactly a formalism can describe a language relative to other formalisms -- is a c…