1 paper
Bingzheng Gan, Tianyi Zhang, Yusu Li +4
The scalability of Large Language Models to long sequences is hindered by the quadratic cost of attention and the limitations of positional encodings. To address these, we introduc…