1 paper · 1 filter
Chen Zhang, Yan Ding, Haotian Wang +3
During the deployment of Large Language Models (LLMs), the autoregressive decoding phase on heterogeneous NPU platforms (e.g., Ascend 910B) faces severe memory-bound challenges. Th…