1 paper
Yuxuan Liu, Wenyuan Li, Laizhong Cui +1
Large language models (LLMs) often face a bottleneck in inference speed due to their reliance on auto-regressive decoding. Recently, parallel decoding has shown significant promise…