Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
Why Does the Effective Context Length of LLMs Fall Short?
Chenxin An, Jun Zhang, Ming Zhong +5
Advancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work r…
cs.CL2024
Training-Free Long-Context Scaling of Large Language Models
Chenxin An, Fei Huang, Jun Zhang +4
The ability of Large Language Models (LLMs) to process and generate coherent text is markedly weakened when the number of input tokens exceeds their pretraining length. Given the e…
cs.CL2023
Linear Attention via Orthogonal Memory
Jun Zhang, Shuyang Jiang, Jiangtao Feng +2
Efficient attentions have greatly improved the computational efficiency of Transformers. However, most existing linear attention mechanisms suffer from an \emph{efficiency degradat…