3 papers
cs.LG2025
Simulating Hard Attention Using Soft Attention
Andy Yang, Lena Strobl, David Chiang +1
We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine…
cs.LG2025
Concise One-Layer Transformers Can Do Function Evaluation (Sometimes)
Lena Strobl, Dana Angluin, Robert Frank
While transformers have proven enormously successful in a range of tasks, their fundamental properties as models of computation are not well understood. This paper contributes to t…
cs.FL2024
Transformers as Transducers
Lena Strobl, Dana Angluin, David Chiang +2
We study the sequence-to-sequence mapping capacity of transformers by relating them to finite transducers, and find that they can express surprisingly large classes of transduction…