From the 1 of 5 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Characterizing the Expressivity of Local Attention in Transformers
Jiaoda Li, Ryan Cotterell
The transformer is the most popular neural architecture for language modeling. The cornerstone of the transformer is its global attention mechanism, which lets the model aggregate…
cs.CL2026
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
Brian DuSell, Ryan Cotterell
When children learn language, they make syntactic generalizations based on hierarchical rules. A recent line of work has inquired as to whether common neural network architectures…