2 papers
cs.AR2025
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
Jinhao Li, Jiaming Xu, Shan Huang +9
Large Language Models (LLMs) have demonstrated remarkable capabilities across various fields, from natural language understanding to text generation. Compared to non-generative LLM…
cs.DC2025
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
Jiaming Xu, Jiayi Pan, Yongkang Zhou +5
Early exiting has recently emerged as a promising technique for accelerating large language models (LLMs) by effectively reducing the hardware computation and memory access. In thi…