1 citations · 1 across the 1 of their papers we have counts for
1 paper
Shidi Li, Christian Walder, Alexander Soen +2
The sparse transformer can reduce the computational complexity of the self-attention layers to O(n), whilst still being a universal approximator of continuous sequence-to-sequenc…