1 paper · 1 filter
Yifan Chen, Qi Zeng, Dilek Hakkani-Tur +3
Transformer-based models are not efficient in processing long sequences due to the quadratic space and time complexity of the self-attention modules. To address this limitation, Li…