49 citations · 80 across the 14 of their papers we have counts for
1 paper · 1 filter
Xuezhe Ma, Xiang Kong, Sinong Wang +4
The quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Lun…