25 citations · 27 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 2 cited
You Only Cache Once: Decoder-Decoder Architectures for Language Models
Yutao Sun, Li Dong, Yi Zhu +6
We introduce a decoder-decoder architecture, YOCO, for large language models, which only caches key-value pairs once. It consists of two components, i.e., a cross-decoder stacked u…
cs.CL2023★ 25 cited
Efficient Large Language Models: A Survey
Zhongwei Wan, Xin Wang, Che Liu +9
Large Language Models (LLMs) have demonstrated remarkable capabilities in important tasks such as natural language understanding and language generation, and thus have the potentia…