8 papers
DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge
Yaodan Xu, Sheng Zhou, Zhisheng Niu
Speculative decoding has emerged as a promising technique for large language model (LLM) inference by accelerating autoregressive decoding via draft-then-verify. This paper studies…
Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference
Yunchu Han, Zhaojun Nan, Sheng Zhou +1
Deep neural networks (DNNs) have been widely applied in diverse applications, but the problems of high latency and energy overhead are inevitable on resource-constrained devices. T…
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
Zhaojun Nan, Yunchu Han, Sheng Zhou +1
In edge intelligence systems, deep neural network (DNN) partitioning and data offloading can provide real-time task inference for resource-constrained mobile devices. However, the…
FedTeddi: Temporal Drift and Divergence Aware Scheduling for Timely Federated Edge Learning
Yuxuan Bai, Yuxuan Sun, Tan Chen +3
Federated edge learning (FEEL) enables collaborative model training across distributed clients over wireless networks without exposing raw data. While most existing studies assume…
DVFS-Aware DNN Inference on GPUs: Latency Modeling and Performance Analysis
Yunchu Han, Zhaojun Nan, Sheng Zhou +1
The rapid development of deep neural networks (DNNs) is inherently accompanied by the problem of high computational costs. To tackle this challenge, dynamic voltage frequency scali…
FedCGD: Collective Gradient Divergence Optimized Scheduling for Wireless Federated Learning
Tan Chen, Jintao Yan, Yuxuan Sun +2
Federated learning (FL) is a promising paradigm for multiple devices to cooperatively train a model. When applied in wireless networks, two issues consistently affect the performan…