20 papers
A Sovereign, Open-Source Foundation Model for German and English
The Soofi-Team, Soofi-Team, : +31
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B…
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
Rupert Mitchell, Kristian Kersting
Pretraining transformers on long sequences (entire code repositories, collections of related documents) is bottlenecked by quadratic attention costs. We present Multipole Semantic…
COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives
David Steinmann, Antonia Wüst, Kristian Kersting +1
While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to s…
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?
Nils Grandien, David Steinmann, Felix Friedrich +1
Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable,…
Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation
Selen Erkan, Bastian Boll, Kristian Kersting +2
Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific formatting requirements. This esp…
KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
Maurice Kraus, Ruben Härle, Sebastian Sztwiertnia +5
High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are o…