15 citations · 29 across the 10 of their papers we have counts for
Showing 2019 · cs.CLShow all
2 papers · 2 filters
cs.CL2019
Adaptively Sparse Transformers
Gonçalo M. Correia, Vlad Niculae, André F. T. Martins
Attention mechanisms have become ubiquitous in NLP. Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed a…
cs.CL2019★ 7 cited
Sparse Sequence-to-Sequence Models
Ben Peters, Vlad Niculae, André F. T. Martins
Sequence-to-sequence models are a powerful workhorse of NLP. Most variants employ a softmax transformation in both their attention mechanism and output layer, leading to dense alig…