8 papers
Decentralized Rank Scheduling for Energy-Constrained Multi-Task Federated Fine-Tuning in Edge-Assisted IoV Networks
Bokeng Zheng, Jianqiang Zhong, Jiayi Liu +3
Large-scale Internet of Vehicles (IoV) deployments increasingly demand the on-device adaptation of foundation models to support diverse, mission-critical perception tasks. While fe…
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
Tian Wu, Liming Wang, Zijian Wen +5
The emergence of Mixture-of-Experts (MoE) has transformed the scaling of large language models by enabling vast model capacity through sparse activation. Yet, converting these perf…
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
Liangkun Chen, Zijian Wen, Tian Wu +2
The Mixture-of-Experts (MoE) architecture has been widely adopted in large language models (LLMs) to reduce computation cost through model sparsity. Employing speculative decoding…
Spatio-Temporal Parallelism for Diffusion Model Inference on Heterogeneous Multi-GPU Systems
Han Liang, Jiahui Zhou, Zicheng Zhou +2
The widespread adoption of diffusion models for image generation necessitates efficient parallel inference to manage their substantial computational overhead. However, current para…
TACO: Tackling Over-correction in Federated Learning with Tailored Adaptive Correction
Weijie Liu, Ziwei Zhan, Carlee Joe-Wong +5
Non-independent and identically distributed (Non-IID) data across edge clients have long posed significant challenges to federated learning (FL) training in edge computing environm…
Real-Time Neural-Enhancement for Online Cloud Gaming
Shan Jiang, Zhenhua Han, Haisheng Tan +6
Online Cloud gaming demands real-time, high-quality video transmission across variable wide-area networks (WANs). Neural-enhanced video transmission algorithms employing super-reso…