10 citations · 46 across the 10 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
Zhenyu Zhang, Runjin Chen, Shiwei Liu +5
This paper aims to overcome the "lost-in-the-middle" challenge of large language models (LLMs). While recent advancements have successfully enabled LLMs to perform stable language…
cs.CL2023
ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks
Xiaoxia Wu, Haojun Xia, Stephen Youn +9
This study examines 4-bit quantization methods like GPTQ in large language models (LLMs), highlighting GPTQ's overfitting and limited enhancement in Zero-Shot tasks. While prior wo…
cs.CL2022★ 4 cited
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers
Zhewei Yao, Xiaoxia Wu, Conglong Li +4
Large-scale transformer models have become the de-facto architectures for various machine learning applications, e.g., CV and NLP. However, those large models also introduce prohib…