7 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.DC2026
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
Zizhao Mo, Junlin Chen, Huanle Xu +1
Nowadays, service providers often deploy multiple types of LLM services within shared clusters. While the service colocation improves resource utilization, it introduces significan…
cs.DC2022★ 7 cited
PECCO: A Profit and Cost-oriented Computation Offloading Scheme in Edge-Cloud Environment with Improved Moth-flame Optimisation
Jiashu Wu, Hao Dai, Yang Wang +2
With the fast growing quantity of data generated by smart devices and the exponential surge of processing demand in the Internet of Things (IoT) era, the resource-rich cloud centre…
cs.DS2022★ 5 cited
PackCache: An Online Cost-driven Data Caching Algorithm in the Cloud
Jiashu Wu, Hao Dai, Yang Wang +3
In this paper, we study a data caching problem in the cloud environment, where multiple frequently co-utilised data items could be packed as a single item being transferred to serv…