8 citations · 16 across the 48 of their papers we have counts for
19 papers · 1 filter
Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations
Antonin Poché, Fanny Jourdan, Nils Feldhus +6
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is co…
Compositional Multilingual and Behavioral Attribute Steering
Hyun Gu Kang, Daniil Gurgurov, Tanja Baeumel +2
This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we inve…
When Tokenization is Secretly Output Supervision
Tanja Baeumel, Josef van Genabith, Simon Ostermann
Tokenization in language models is treated by default as an input preprocessing decision. We argue that this framing is incomplete: in autoregressive models, tokenizer granularity…
Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov +2
Multilingual large language models (mLLMs) achieve strong performance in machine translation, yet our understanding of the mechanisms by which they transform representations from o…
A Sovereign, Open-Source Foundation Model for German and English
Soofi-Team, :, Benedikt Droste +30
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B…
Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation
Jan Cegin, Daniil Gurgurov, Yusser Al Ghussin +1
Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated data can improve downstream task per…