1 paper · 1 filter
Ashkan Shahbazi, Chayne Thrash, Yikun Bai +3
Transformers have proven highly effective across modalities, but standard softmax attention scales quadratically with sequence length, limiting long context modeling. Linear attent…