2 citations · 2 across the 2 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Simulating Hard Attention Using Soft Attention
Andy Yang, Lena Strobl, David Chiang +1
We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine…
cs.FL2024
Transformers as Transducers
Lena Strobl, Dana Angluin, David Chiang +2
We study the sequence-to-sequence mapping capacity of transformers by relating them to finite transducers, and find that they can express surprisingly large classes of transduction…