7 papers
Reliable Microservice Tail Latency Prediction via Decoupled Dual-Stream Learning and Gradient Modulation
Wenzhuo Qian, Hailiang Zhao, Jiayi Chen +7
Microservice architectures enable scalable cloud-native applications; however, the distributed nature of these systems complicates the maintenance of strict Service Level Objective…
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
Peng Chen, Jiaji Zhang, Hailiang Zhao +11
In modern GPU inference, cache efficiency remains a major bottleneck, and heuristic policies such as \textsc{LRU} can perform far worse than the offline optimum. Existing learning-…
SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
Jiaji Zhang, Ruichao Sun, Hailiang Zhao +7
Diffusion models have demonstrated exceptional generative capabilities but are computationally intensive, posing significant challenges for deployment in resource-constrained or la…
Morphis: SLO-Aware Resource Scheduling for Microservices with Time-Varying Call Graphs
Yu Tang, Hailiang Zhao, Chuansheng Lu +4
Modern microservice systems exhibit continuous structural evolution in their runtime call graphs due to workload fluctuations, fault responses, and deployment activities. Despite t…
Adaptive Dual-Weighting Framework for Federated Learning via Out-of-Distribution Detection
Zhiwei Ling, Hailiang Zhao, Chao Zhang +8
Federated Learning (FL) enables collaborative model training across large-scale distributed service nodes while preserving data privacy, making it a cornerstone of intelligent serv…
CLMTracing: Black-box User-level Watermarking for Code Language Model Tracing
Boyu Zhang, Ping He, Tianyu Du +4
With the widespread adoption of open-source code language models (code LMs), intellectual property (IP) protection has become an increasingly critical concern. While current waterm…