12 papers
Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training
Niklas Enskat, Philipp Wiesner
High-fidelity performance simulators are essential for designing and configuring efficient AI systems, yet today's tools lack the ability to predict power consumption. Established…
Exploring Silent Data Corruption as a Reliability Challenge in LLM Training
Anton Altenbernd, Philipp Wiesner, Odej Kao
As Large Language Models (LLMs) scale in size and complexity, the consequences of failures during training become increasingly severe. A major challenge arises from Silent Data Cor…
Carbon-Aware Quality Adaptation for Energy-Intensive Services
Philipp Wiesner, Dennis Grinwald, Philipp Weià +3
The energy demand of modern cloud services, particularly those related to generative AI, is increasing at an unprecedented pace. To date, carbon-aware computing strategies have pri…
Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study
Philipp Wiesner, Soeren Becker, Brett Cornick +3
Training large language models (LLMs) requires substantial compute and energy. At the same time, renewable energy sources regularly produce more electricity than the grid can absor…
Efficiency Will Not Lead to Sustainable Reasoning AI
Philipp Wiesner, Daniel W. O'Neill, Francesca Larosa +1
AI research is increasingly moving toward complex problem solving, where models are optimized not only for pattern recognition but for multi-step reasoning. Historically, computing…
What happens when nanochat meets DiLoCo?
Alexander Acker, Soeren Becker, Sasho Nedelkoski +3
Although LLM training is typically centralized with high-bandwidth interconnects and large compute budgets, emerging methods target communication-constrained training in distribute…