1 citations · 1 across the 17 of their papers we have counts for
5 papers · 1 filter
KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
Maurice Kraus, Ruben Härle, Sebastian Sztwiertnia +5
High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are o…
A Sovereign, Open-Source Foundation Model for German and English
Soofi-Team, :, Benedikt Droste +30
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B…
Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation
Selen Erkan, Bastian Boll, Kristian Kersting +2
Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific formatting requirements. This esp…
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
Max Henning Höth, Kristian Kersting, Björn Deiseroth +1
Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both contributes to and faithfully…
Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models
Björn Deiseroth, Max Henning Höth, Kristian Kersting +1
Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic rather than verifiable evidence -- le…