1 paper · 1 filter
Ruize He, Dongchen Han, Gao Huang
Existing research largely attributes the global sequence modeling capability of Transformers to the explicit computation of attention weights, a process that inherently incurs quad…