2.3k citations · 2.9k across the 9 of their papers we have counts for
1 paper · 1 filter
Yufan Jiang, Qiaozhi He, Xiaomin Zhuang +4
Existing large language models have to run K times to generate a sequence of K tokens. In this paper, we present RecycleGPT, a generative language model with fast decoding speed by…