5 papers
Kirin: Improving ANN efficiency with SNN Hybridization
Chenyu Wang, Zhanglu Yan, Zhi Zhou +2
Artificial neural networks (ANNs), particularly large language models (LLMs), demonstrate powerful inference capabilities but consume substantial energy. Conversely, spiking neural…
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
Guihang Hong, Tao Ouyang, Kongyange Zhao +2
Motivated by the imperative for real-time responsiveness and data privacy preservation, large language models (LLMs) are increasingly deployed on resource-constrained edge devices…
Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
Chenyu Wang, Zhanglu Yan, Zhi Zhou +2
In the era of large language models (LLMs), weight-activation quantization helps fit models on edge device by reducing memory and compute bit-widths. However, three challenges pers…
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
Jin Yang, Qiong Wu, Zhiying Feng +3
Large Language Models (LLMs) have demonstrated remarkable capabilities, leading to a significant increase in user demand for LLM services. However, cloud-based LLM services often s…
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
Yunkai Liang, Zhangyu Chen, Pengfei Zuo +3
In large language model (LLM) serving systems, executing each request consists of two phases: the compute-intensive prefill phase and the memory-intensive decoding phase. To preven…