5 papers
Knee-Deep in C-RASP: A Transformer Depth Hierarchy
Andy Yang, Michaël Cadilhac, David Chiang
It has been observed that transformers with greater depth (that is, more layers) have more capabilities, but can we establish formally which capabilities are gained? We answer this…
Simulating Hard Attention Using Soft Attention
Andy Yang, Lena Strobl, David Chiang +1
We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine…
Counting Like Transformers: Compiling Temporal Counting Logic Into Softmax Transformers
Andy Yang, David Chiang
Deriving formal bounds on the expressivity of transformers, as well as studying transformers that are constructed to implement known algorithms, are both effective methods for bett…
Transformers as Transducers
Lena Strobl, Dana Angluin, David Chiang +2
We study the sequence-to-sequence mapping capacity of transformers by relating them to finite transducers, and find that they can express surprisingly large classes of transduction…
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
Andy Yang, David Chiang, Dana Angluin
The expressive power of transformers over inputs of unbounded size can be studied through their ability to recognize classes of formal languages. In this paper, we establish exact…