1 paper · 1 filter
Feng Lin, Hanling Yi, Hongbin Li +4
Large language models (LLMs) commonly employ autoregressive generation during inference, leading to high memory bandwidth demand and consequently extended latency. To mitigate this…