3 papers
cs.LG2025
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
Mogens Henrik From, Jacob Nielsen, Lukas Galke Poech +1
Training large neural network models requires extensive computational resources, often distributed across several nodes and accelerators. Recent findings suggest that it may be suf…
cs.LG2025
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
Jacob Nielsen, Peter Schneider-Kamp, Lukas Galke
Large language models (LLMs) require immense resources for training and inference. Quantization, a technique that reduces the precision of model parameters, offers a promising solu…
cs.LG2024
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
Jacob Nielsen, Lukas Galke, Peter Schneider-Kamp
Contemporary machine learning models, such as language models, are powerful, but come with immense resource requirements both at training and inference time. It has been shown that…