collaborators

6 papers

cs.DC2026

Towards Energy Efficient Co-Scheduling in HPC

Zhong Zheng, Michael E. Papka, Zhiling Lan

Modern multi GPU HPC systems expose substantial computational capacity, yet inefficient GPU allocation often leads to wasted energy and underutilization. In practice, GPU applicati…

cs.DC2026

EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems

Zhong Zheng, Michael E. Papka, Zhiling Lan

Power-constrained HPC systems increasingly run heterogeneous CPU--GPU applications under strict cluster-wide power limits. Existing cluster-wide power management policies rely on f…

cs.DC2026

Understanding Large-Scale HPC System Behavior Through Cluster-Based Visual Analytics

Allison Austin, Shilpika, Yan To Linus Lam +4

In high-performance computing (HPC) environments, system monitoring data is often unlabeled and high-dimensional, making it difficult to reliably detect and understand anomalous co…

cs.DC2025

A Real-Time Digital Twin for Adaptive Scheduling

Yihe Zhang, Yash Kurkure, Yiheng Tao +2

High-performance computing (HPC) workloads are becoming increasingly diverse, exhibiting wide variability in job characteristics, yet cluster scheduling has long relied on static,…

cs.DC2025

Understanding the Landscape of Ampere GPU Memory Errors

Zhu Zhu, Yu Sun, Dhatri Parakal +9

Graphics Processing Units (GPUs) have become a de facto solution for accelerating high-performance computing (HPC) applications. Understanding their memory error behavior is an ess…

cs.DC2025

Coordinated Power Management on Heterogeneous Systems

Zhong Zheng, Zhiling Lan, Xingfu Wu +2

Performance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling method…