1 paper
Peipei Li, Dongsen Zhang, Yuchen Liu +1
Large Language Models (LLMs) primarily perform inference at the token level, resulting in substantial memory overhead and compromised computational efficiency. In this paper, we pr…