1 paper
Ziwei He, Meng Yang, Minwei Feng +4
The transformer model is known to be computationally demanding, and prohibitively costly for long sequences, as the self-attention module uses a quadratic time and space complexity…