1 paper
Mingbo Hao, Changwei Yan, Haoyu Cui +4
The rapid growth of LLMs demands high-throughput, memory-capacity-intensive inference on resource-constrained edge devices, where single-batch decoding remains fundamentally memory…