Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge
Tianhao Jiang, Hang Gu, Teng Wang +9
Edge LLM inference combines sparsity and low-bit quantization to meet device memory, latency, and power limits. Yet quantization shrinks weight payloads without proportionally redu…
cs.DC2026
Accelerating OpenPangu Inference on NPU via Speculative Decoding
Yuntao Dai, Jing Wu, Hang Gu +1
To mitigate the Memory Wall bottleneck encountered by Large Language Models (LLMs) during inference on \textbf{NPU} hardware, and addressing the scarcity of native support for main…