4 citations · 4 across the 12 of their papers we have counts for
1 paper · 1 filter
Zeping Li, Xinlong Yang, Ziheng Gao +7
Large Language Models (LLMs) inherently use autoregressive decoding, which lacks parallelism in inference and results in significantly slow inference speed. While methods such as M…