2 papers
cs.AR2026
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
Mingbo Hao, Changwei Yan, Haoyu Cui +4
The rapid growth of LLMs demands high-throughput, memory-capacity-intensive inference on resource-constrained edge devices, where single-batch decoding remains fundamentally memory…
cs.AR2025
KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
Lishuo Deng, Shaojie Xu, Jinwu Chen +4
Deploying large language models (LLMs) on edge devices enables personalized agents with strong privacy and low cost. However, with tens to hundreds of billions of parameters, singl…