8 papers
MARS: A Monte Carlo Tree Search-based Adaptive and Responsive Scheduler
Yash Kurkure, Yihe Zhang, Zhiling Lan +1
Modern High Performance Computing systems depend on static heuristics and manual administration for job scheduling and reservation management. Deep Reinforcement Learning (DRL) has…
EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving
Vittorio Palladino, Gianluca Palermo, Michael E. Papka +1
As large language models span dense, mixture-of-experts, and state-space architectures and are deployed on heterogeneous accelerators under increasingly diverse multimodal workload…
Coordinated Power Management on Heterogeneous Systems
Zhong Zheng, Zhiling Lan, Xingfu Wu +2
Performance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling method…
Towards Energy Efficient Co-Scheduling in HPC
Zhong Zheng, Michael E. Papka, Zhiling Lan
Modern multi GPU HPC systems expose substantial computational capacity, yet inefficient GPU allocation often leads to wasted energy and underutilization. In practice, GPU applicati…
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
Zhong Zheng, Michael E. Papka, Zhiling Lan
Power-constrained HPC systems increasingly run heterogeneous CPU--GPU applications under strict cluster-wide power limits. Existing cluster-wide power management policies rely on f…
A Real-Time Digital Twin for Adaptive Scheduling
Yihe Zhang, Yash Kurkure, Yiheng Tao +2
High-performance computing (HPC) workloads are becoming increasingly diverse, exhibiting wide variability in job characteristics, yet cluster scheduling has long relied on static,…