2 papers
cs.AR2025
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
Alessandro Fogli, Bo Zhao, Peter Pietzuch +1
The growing disparity between CPU core counts and available memory bandwidth has intensified memory contention in servers. This particularly affects highly parallelizable applicati…
cs.DC2024
Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
Marcel Wagenländer, Guo Li, Bo Zhao +2
Deep learning (DL) jobs use multi-dimensional parallelism, i.e. combining data, model, and pipeline parallelism, to use large GPU clusters efficiently. Long-running jobs may experi…