From the 1 of 6 linked papers with an AI index.
1 paper · 1 filter
Tongxi Wang
Large language models (LLMs) excel across many tasks, yet inference is still dominated by strictly token-by-token autoregression. Existing acceleration methods largely patch this p…