4 papers
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors
Alexander Hägele, Alejandro Hernández-Cano, Atli Kosson +1
Modern neural network training relies on optimizers such as Adam and Muon which act on each weight matrix as a single object. Yet every weight matrix carries two distinct quantitie…
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
Jonathan Coles, Stefano Schuppli, Lukas Drescher +20
Large Language Models (LLMs) have surged as a transformative technology for science and society, prompting governments worldwide to pursue sovereign AI capabilities that ensure dat…
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100
We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…
Towards Fully FP8 GEMM LLM Training at Scale
Alejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag +1
Despite the significant potential of FP8 data formats for large language model (LLM) pre-training, their adoption has been limited due to challenges in maintaining stability at sca…