1 paper · 1 filter
Aishwarya P S, Pranav Ajit Nair, Yashas Samaga +4
The autoregressive nature of conventional large language models (LLMs) inherently limits inference speed, as tokens are generated sequentially. While speculative and parallel decod…