3 papers
cs.PF2026
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
Kaixuan Zhang, Chutong Ding, Shiyou Qian +6
The rapid adoption of Large Language Models (LLMs) has made GPU inference efficiency an increasingly critical system concern. The runtime of LLM workloads is largely dominated by t…
cs.PF2026
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
Kaixuan Zhang, Yunfan Cui, Shuhao Zhang +8
The rapid expansion of Transformer-based large language models has dramatically increased the need for high-performance GPUs. As a result, there is growing demand for fast, accurat…
cs.DC2024
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
Dingyu Yang, Kangpeng Zheng, Shiyou Qian +2
Co-locating latency-critical services (LCSs) and best-effort jobs (BEJs) constitute the principal approach for enhancing resource utilization in production. Nevertheless, the co-lo…