1 paper
Seongjun Yang, Gibbeum Lee, Jaewoong Cho +2
This paper presents "Predictive Pipelined Decoding (PPD)," an approach that speeds up greedy decoding in Large Language Models (LLMs) while maintaining the exact same output as the…