3 papers
cs.DC2025
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
Yi Pan, Wenbo Qian, Dedong Xie +3
The training and deployment of machine learning (ML) models have become extremely energy-intensive. While existing optimization efforts focus primarily on hardware energy efficienc…
cs.PF2025
gigiProfiler: Diagnosing Performance Issues by Uncovering Application Resource Bottlenecks
Yigong Hu, Haodong Zheng, Yicheng Liu +3
Diagnosing performance bottlenecks in modern software is essential yet challenging, particularly as applications become more complex and rely on custom resource management policies…
cs.DC2025
NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yufei Gao, Yilong Zhao +13
Large Language Models (LLMs) have resulted in a surging demand for planet-scale serving systems, where tens of thousands of GPUs continuously serve hundreds of millions of users. C…