1 paper
Bin Xiao, Lujun Gui, Lei Su +1
Large Language Models (LLMs) frequently suffer from inefficiencies, largely attributable to the discord between the requirements of auto-regressive decoding and the architecture of…