1 paper · 1 filter
Lexington Whalen, Yuki Ito, Ryo Sakamoto
Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inhere…