1 citations · 1 across the 21 of their papers we have counts for
1 paper · 2 filters
Zeping Li, Xinlong Yang, Ziheng Gao +7
Large Language Models (LLMs) inherently use autoregressive decoding, which lacks parallelism in inference and results in significantly slow inference speed. While methods such as M…