Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
Carolin Penke, Chelsea Maria John, Jan Ebert +2
The training of large language models (LLMs) requires substantial computational resources, complex software stacks, and carefully designed workflows to achieve scalability and effi…
cs.DC2025
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
Jiangtao Wang, Jan Ebert, Oleg Filatov +1
Transformer models have revolutionized a wide spectrum of disciplines, especially in language processing. The recent success has proven that model size scalability is crucial for a…