3 papers
cs.AR2025
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
Tianhua Xia, Sai Qian Zhang
Running Large Language Models (LLMs) on edge devices is crucial for reducing latency, improving real-time processing, and enhancing privacy. By performing inference directly on the…
cs.AR2025
HAAN: A Holistic Approach for Accelerating Normalization Operations in Large Language Models
Tianfan Peng, Jiajun Qin, Tianhua Xia +1
Large language models (LLMs) have revolutionized natural language processing (NLP) tasks by achieving state-of-the-art performance across a range of benchmarks. Central to the succ…
cs.LG2025
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Patrick Yubeaton, Tareq Mahmoud, Shehab Naga +8
As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small,…