9 papers
Preserving Admission Responsibility in Multi-Tenant Large Language Model Prefix Caches
Zhiyu Wang, Rajkumar Buyya
Shared prefix caching turns Graphics Processing Unit (GPU) memory into persistent state shared across Large Language Model (LLM) tenants. A group that materializes new Key-Value (K…
PrefixPlace: Provable Prefix Key-Value Placement for Large Language Model Serving under Heterogeneous Compute and Transfer Costs
Zhiyu Wang, Rajkumar Buyya
Prefix Key-Value (KV) reuse avoids repeated prefill in Large Language Model (LLM) inference, but local misses require recomputation or replica fetches. Their relative cost varies w…
ReinFog: A Deep Reinforcement Learning Empowered Framework for Resource Management in Edge and Cloud Computing Environments
Zhiyu Wang, Mohammad Goudarzi, Rajkumar Buyya
The growing IoT landscape requires effective server deployment strategies to meet demands including real-time processing and energy efficiency. This is complicated by heterogeneous…
DeFRiS: Silo-Cooperative IoT Applications Scheduling via Decentralized Federated Reinforcement Learning
Zhiyu Wang, Mohammad Goudarzi, Mingming Gong +1
Next-generation IoT applications increasingly span across autonomous administrative entities, necessitating silo-cooperative scheduling to leverage diverse computational resources…
A Risk-Aware UAV-Edge Service Framework for Wildfire Monitoring and Emergency Response
Yulun Huang, Zhiyu Wang, Rajkumar Buyya
Wildfire monitoring demands timely data collection and processing for early detection and rapid response. UAV-assisted edge computing is a promising approach, but jointly minimizin…
Performance and Security Aware Distributed Service Placement in Fog Computing
Mohammad Goudarzi, Arash Shaghaghi, Zhiyu Wang +1
The rapid proliferation of IoT applications has intensified the demand for efficient and secure service placement in Fog computing. However, heterogeneous resources, dynamic worklo…