3 papers
cs.MA2026
Safe Multi-Agent Deep Reinforcement Learning for Privacy-Aware Edge-Device Collaborative DNN Inference
Hong Wang, Xuwei Fan, Zhipeng Cheng +4
As Deep Neural Network (DNN) inference becomes increasingly prevalent on edge and mobile platforms, critical challenges emerge in privacy protection, resource constraints, and dyna…
cs.AR2025
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
William Meng, Benjamin Lee, Hong Wang
KV cache offloading enables long-context LLM inference by storing caches in CPU DRAM, but PCIe bandwidth limitations create severe bottlenecks. In this paper, we develops an analyt…
cs.LG2025
Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services
Zhipeng Cheng, Xiaoyu Xia, Hong Wang +4
Edge inference (EI) has emerged as a promising paradigm to address the growing limitations of cloud-based Deep Neural Network (DNN) inference services, such as high response latenc…