22 citations · 30 across the 5 of their papers we have counts for
Showing 2021Show all
2 papers · 1 filter
cs.CL2021★ 1 cited
Regularized Training of Nearest Neighbor Language Models
Jean-Francois Ton, Walter Talbott, Shuangfei Zhai +1
Including memory banks in a natural language processing architecture increases model capacity by equipping it with additional data at inference time. In this paper, we build upon $…
cs.LG2021
An Attention Free Transformer
Shuangfei Zhai, Walter Talbott, Nitish Srivastava +4
We introduce Attention Free Transformer (AFT), an efficient variant of Transformers that eliminates the need for dot product self attention. In an AFT layer, the key and value are…