6 papers
Kirin: Improving ANN efficiency with SNN Hybridization
Chenyu Wang, Zhanglu Yan, Zhi Zhou +2
Artificial neural networks (ANNs), particularly large language models (LLMs), demonstrate powerful inference capabilities but consume substantial energy. Conversely, spiking neural…
Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
Chenyu Wang, Zhanglu Yan, Zhi Zhou +2
In the era of large language models (LLMs), weight-activation quantization helps fit models on edge device by reducing memory and compute bit-widths. However, three challenges pers…
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
Jin Yang, Qiong Wu, Zhiying Feng +3
Large Language Models (LLMs) have demonstrated remarkable capabilities, leading to a significant increase in user demand for LLM services. However, cloud-based LLM services often s…
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
Yunkai Liang, Zhangyu Chen, Pengfei Zuo +3
In large language model (LLM) serving systems, executing each request consists of two phases: the compute-intensive prefill phase and the memory-intensive decoding phase. To preven…
Online Location Planning for AI-Defined Vehicles: Optimizing Joint Tasks of Order Serving and Spatio-Temporal Heterogeneous Model Fine-Tuning
Bokeng Zheng, Bo Rao, Tianxiang Zhu +5
Advances in artificial intelligence (AI) including foundation models (FMs), are increasingly transforming human society, with smart city driving the evolution of urban living.Meanw…
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
Rui Li, Tao Ouyang, Liekang Zeng +3
Collaborative Edge Computing (CEC) is an emerging paradigm that collaborates heterogeneous edge devices as a resource pool to compute DNN inference tasks in proximity such as edge…