25 citations · 65 across the 5 of their papers we have counts for
5 papers
Union: An Automatic Workload Manager for Accelerating Network Simulation
Xin Wang, Misbah Mubarak, Yao Kang +2
With the rapid growth of the machine learning applications, the workloads of future HPC systems are anticipated to be a mix of scientific simulation, big data analytics, and machin…
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
Yao Kang, Xin Wang, Zhiling Lan
High-radix interconnects such as Dragonfly and its variants rely on adaptive routing to balance network traffic for optimum performance. Ideally, adaptive routing attempts to forwa…
MRSch: Multi-Resource Scheduling for HPC
Boyang Li, Yuping Fan, Matthew Dearing +4
Emerging workloads in high-performance computing (HPC) are embracing significant changes, such as having diverse resource requirements instead of being CPU-centric. This advancemen…
Study of Workload Interference with Intelligent Routing on Dragonfly
Yao Kang, Xin Wang, Zhiling Lan
Dragonfly interconnect is a crucial network technology for supercomputers. To support exascale systems, network resources are shared such that links and routers are not dedicated t…
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
Boyang Li, Zhiling Lan, Michael E. Papka
In the field of high-performance computing (HPC), there has been recent exploration into the use of deep reinforcement learning for cluster scheduling (DRL scheduling), which has d…