1 paper · 1 filter
Hongjun An, Yifan Chen, Zhe Sun +1
Current large language models (LLMs) primarily utilize next-token prediction method for inference, which significantly impedes their processing speed. In this paper, we introduce a…