6 papers
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
Mogens Henrik From, Jacob Nielsen, Lukas Galke Poech +1
Training large neural network models requires extensive computational resources, often distributed across several nodes and accelerators. Recent findings suggest that it may be suf…
Isolating Culture Neurons in Multilingual Large Language Models
Danial Namazifard, Lukas Galke Poech
Language and culture are deeply intertwined, yet it has been unclear how and where multilingual large language models encode culture. Here, we build on an established methodology f…
Guarded Query Routing for Large Language Models
Richard Šléher, William Brach, Tibor Sloboda +2
Query routing, the task to route user queries to different large language model (LLM) endpoints, can be considered as a text classification problem. However, out-of-distribution qu…
Super-additive Cooperation in Language Model Agents
Filippo Tonini, Lukas Galke
With the prospect of autonomous artificial intelligence (AI) agents, studying their tendency for cooperative behavior becomes an increasingly relevant topic. This study is inspired…
Dynaword: From One-shot to Continuously Developed Datasets
Kenneth Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan +14
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguousl…
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
Jacob Nielsen, Peter Schneider-Kamp, Lukas Galke
Large language models (LLMs) require immense resources for training and inference. Quantization, a technique that reduces the precision of model parameters, offers a promising solu…