1 paper
Jinhyeok Kim, Yejoon Lee, Jaeyoung Do
The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwid…