1 paper
Yang Zhou, Zhuoming Chen, Zhaozhuo Xu +2
With the blossom of large language models (LLMs), inference efficiency becomes increasingly important. Various approximation methods are proposed to reduce the cost at inference ti…