collaborators

6 papers

cs.LG2025

DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes

Mogens Henrik From, Jacob Nielsen, Lukas Galke Poech +1

Training large neural network models requires extensive computational resources, often distributed across several nodes and accelerators. Recent findings suggest that it may be suf…

cs.CL2025

Isolating Culture Neurons in Multilingual Large Language Models

Danial Namazifard, Lukas Galke Poech

Language and culture are deeply intertwined, yet it has been unclear how and where multilingual large language models encode culture. Here, we build on an established methodology f…

cs.AI2025

Guarded Query Routing for Large Language Models

Richard Šléher, William Brach, Tibor Sloboda +2

Query routing, the task to route user queries to different large language model (LLM) endpoints, can be considered as a text classification problem. However, out-of-distribution qu…

cs.AI2025

Super-additive Cooperation in Language Model Agents

Filippo Tonini, Lukas Galke

With the prospect of autonomous artificial intelligence (AI) agents, studying their tendency for cooperative behavior becomes an increasingly relevant topic. This study is inspired…

cs.CL2025

Dynaword: From One-shot to Continuously Developed Datasets

Kenneth Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan +14

Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguousl…

cs.LG2025

Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?

Jacob Nielsen, Peter Schneider-Kamp, Lukas Galke

Large language models (LLMs) require immense resources for training and inference. Quantization, a technique that reduces the precision of model parameters, offers a promising solu…