3 papers
cs.DC2025
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
Ying Zhu, Yang Xu, Hongli Xu +3
Training large language models (LLMs) requires massive computational resources, often necessitating the aggregation of geographically distributed data centers (\ie, cross-region tr…
cs.DC2024
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
Luyao Gao, Jianchun Liu, Hongli Xu +3
End-cloud collaboration offers a promising strategy to enhance the Quality of Service (QoS) in DNN inference by offloading portions of the inference workload from end devices to cl…
cs.DC2024
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
Wenyi Liang, Jianchun Liu, Hongli Xu +2
Edge inference is a technology that enables real-time data processing and analysis on clients near the data source. To ensure compliance with the Service-Level Objectives (SLOs), s…