1 paper · 1 filter
Bin Xiao, Lujun Gui, Lei Su +1
Large Language Models (LLMs) frequently suffer from inefficiencies, largely attributable to the discord between the requirements of auto-regressive decoding and the architecture of…