1 paper · 1 filter
Minsik Cho, Mohammad Rastegari, Devang Naik
Large Language Model or LLM inference has two phases, the prompt (or prefill) phase to output the first token and the extension (or decoding) phase to the generate subsequent token…