1 paper
Martin Benfeghoul, Teresa Delgado, Adnan Oomerjee +3
Transformers' quadratic computational complexity limits their scalability despite remarkable performance. While linear attention reduces this to linear complexity, pre-training suc…