1 paper · 1 filter
Bingzheng Gan, Tianyi Zhang, Yusu Li +4
The scalability of Large Language Models to long sequences is hindered by the quadratic cost of attention and the limitations of positional encodings. To address these, we introduc…