4 papers
What happens when nanochat meets DiLoCo?
Alexander Acker, Soeren Becker, Sasho Nedelkoski +3
Although LLM training is typically centralized with high-bandwidth interconnects and large compute budgets, emerging methods target communication-constrained training in distribute…
Distributed Low-Communication Training with Decoupled Momentum Optimization
Sasho Nedelkoski, Alexander Acker, Odej Kao +2
The training of large models demands substantial computational resources, typically available only in data centers with high-bandwidth interconnects. However, reducing the reliance…
Predicting the Performance of Scientific Workflow Tasks for Cluster Resource Management: An Overview of the State of the Art
Jonathan Bader, Kathleen West, Soeren Becker +5
Scientific workflow management systems support large-scale data analysis on cluster infrastructures. For this, they interact with resource managers which schedule workflow tasks on…
WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows
Fabian Lehmann, Jonathan Bader, Friedrich Tschirpke +6
Scientific workflows process extensive data sets over clusters of independent nodes, which requires a complex stack of infrastructure components, especially a resource manager (RM)…