Showing 2026Show all
2 papers · 1 filter
cs.DC2026
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
Carolin Penke, Chelsea Maria John, Jan Ebert +2
The training of large language models (LLMs) requires substantial computational resources, complex software stacks, and carefully designed workflows to achieve scalability and effi…
cs.LG2026
Optimal Scaling Needs Optimal Norm
Oleg Filatov, Jiangtao Wang, Jan Ebert +1
Despite recent progress in optimal hyperparameter transfer under model and dataset scaling, no unifying explanatory principle has been established. For Adam and Scion optimizers, w…