activity
20242026
collaborators

6 papers

cs.CL2026

A Sovereign, Open-Source Foundation Model for German and English

The Soofi-Team, Soofi-Team, : +31

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B…

cs.CL2026

KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

Maurice Kraus, Ruben Härle, Sebastian Sztwiertnia +5

High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are o…

cs.LG2026

ActivationReasoning: Logical Reasoning in Latent Activation Spaces

Lukas Helff, Ruben Härle, Wolfgang Stammer +6

Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse autoencoders (SAEs) make hidden activatio…

cs.LG2026

LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Lukas Helff, Quentin Delfosse, David Steinmann +6

As reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for scaling reasoning capabilities in LLMs, a new failure mode emerges: LLMs gaming verifi…

cs.CL2025

Measuring and Guiding Monosemanticity

Ruben Härle, Felix Friedrich, Manuel Brack +4

There is growing interest in leveraging mechanistic interpretability and controllability to better understand and influence the internal dynamics of large language models (LLMs). H…

cs.CL2024

SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs

Ruben Härle, Felix Friedrich, Manuel Brack +3

Large Language Models (LLMs) have demonstrated remarkable capabilities in generating human-like text, but their output may not be aligned with the user or even produce harmful cont…