Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Why Does the Effective Context Length of LLMs Fall Short?
Chenxin An, Jun Zhang, Ming Zhong +5
Advancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work r…
cs.CL2024
Training-Free Long-Context Scaling of Large Language Models
Chenxin An, Fei Huang, Jun Zhang +4
The ability of Large Language Models (LLMs) to process and generate coherent text is markedly weakened when the number of input tokens exceeds their pretraining length. Given the e…