collaborators

9 papers

cs.DC2026

MARS: A Monte Carlo Tree Search-based Adaptive and Responsive Scheduler

Yash Kurkure, Yihe Zhang, Zhiling Lan +1

Modern High Performance Computing systems depend on static heuristics and manual administration for job scheduling and reservation management. Deep Reinforcement Learning (DRL) has…

cs.LG2026

Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank

Yiheng Tao, Yihe Zhang, Matthew Dearing +4

Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that is becoming increasingly acute with t…

cs.CV2026

EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

Vittorio Palladino, Gianluca Palermo, Michael E. Papka +1

As large language models span dense, mixture-of-experts, and state-space architectures and are deployed on heterogeneous accelerators under increasingly diverse multimodal workload…

cs.DC2026

Coordinated Power Management on Heterogeneous Systems

Zhong Zheng, Zhiling Lan, Xingfu Wu +2

Performance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling method…

cs.DC2026

Towards Energy Efficient Co-Scheduling in HPC

Zhong Zheng, Michael E. Papka, Zhiling Lan

Modern multi GPU HPC systems expose substantial computational capacity, yet inefficient GPU allocation often leads to wasted energy and underutilization. In practice, GPU applicati…

cs.DC2026

EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems

Zhong Zheng, Michael E. Papka, Zhiling Lan

Power-constrained HPC systems increasingly run heterogeneous CPU--GPU applications under strict cluster-wide power limits. Existing cluster-wide power management policies rely on f…