1 paper
Zicheng Liu, Li Wang, Siyuan Li +3
Transformer models have been successful in various sequence processing tasks, but the self-attention mechanism's computational cost limits its practicality for long sequences. Alth…