1 paper · 1 filter
Ao Sun, Weilin Zhao, Xu Han +4
Effective attention modules have played a crucial role in the success of Transformer-based large language models (LLMs), but the quadratic time and memory complexities of these att…