collaborators

20 papers

cs.CL2026

A Sovereign, Open-Source Foundation Model for German and English

The Soofi-Team, Soofi-Team, : +31

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B…

cs.LG2026

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

Rupert Mitchell, Kristian Kersting

Pretraining transformers on long sequences (entire code repositories, collections of related documents) is bottlenecked by quadratic attention costs. We present Multipole Semantic…

cs.LG2026

COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives

David Steinmann, Antonia Wüst, Kristian Kersting +1

While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to s…

cs.LG2026

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

Nils Grandien, David Steinmann, Felix Friedrich +1

Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable,…

cs.CL2026

Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation

Selen Erkan, Bastian Boll, Kristian Kersting +2

Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific formatting requirements. This esp…

cs.CL2026

KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

Maurice Kraus, Ruben Härle, Sebastian Sztwiertnia +5

High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are o…