7 citations · 14 across the 14 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024★ 1 cited
Self-Selected Attention Span for Accelerating Large Language Model Inference
Tian Jin, Wanzin Yazar, Zifei Xu +2
Large language models (LLMs) can solve challenging tasks. However, their inference computation on modern GPUs is highly inefficient due to the increasing number of tokens they must…
cs.CL2024★ 2 cited
TeleChat Technical Report
Zhongjiang He, Zihan Wang, Xinzhang Liu +33
In this technical report, we present TeleChat, a collection of large language models (LLMs) with parameters of 3 billion, 7 billion and 12 billion. It includes pretrained language…