1 paper
Lexington Whalen, Yuki Ito, Ryo Sakamoto
Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inhere…