13 citations · 19 across the 3 of their papers we have counts for
3 papers
cs.CL2024
On The Adaptation of Unlimiformer for Decoder-Only Transformers
Kian Ahrabian, Alon Benhaim, Barun Patra +3
One of the prominent issues stifling the current generation of large language models is their limited context length. Recent proprietary models such as GPT-4 and Claude 2 have intr…
cs.LG2022★ 6 cited
TorchScale: Transformers at Scale
Shuming Ma, Hongyu Wang, Shaohan Huang +8
Large Transformers have achieved state-of-the-art performance across many tasks. Most open-source libraries on scaling Transformers focus on improving training or inference with be…
cs.LG2022★ 13 cited
Foundation Transformers
Hongyu Wang, Shuming Ma, Shaohan Huang +12
A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different imp…