12 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.AR2025★ 1 cited
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
Alessandro Fogli, Bo Zhao, Peter Pietzuch +1
The growing disparity between CPU core counts and available memory bandwidth has intensified memory contention in servers. This particularly affects highly parallelizable applicati…
cs.DC2023★ 12 cited
Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
Marcel Wagenländer, Guo Li, Bo Zhao +2
Deep learning (DL) jobs use multi-dimensional parallelism, i.e. combining data, model, and pipeline parallelism, to use large GPU clusters efficiently. Long-running jobs may experi…