2 papers
cs.DC2026
ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork
Tim Beringer, Patrick Diem, Felix Wolf +1
Training large-scale AI models often outgrows a single data center, demanding sharded, multi-cluster, and decentralized training. However, the huge space of resource allocations ma…
cs.PF2025
Denoising Application Performance Models with Noise-Resilient Priors
Gustavo de Morais, Alexander Geiß, Alexandru Calotoiu +5
As parallel codes are scaled to larger computing systems, performance models play a crucial role in identifying potential bottlenecks. However, constructing these models analytical…