4 papers · 1 filter
Compositional Multilingual and Behavioral Attribute Steering
Hyun Gu Kang, Daniil Gurgurov, Tanja Baeumel +2
This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we inve…
Tracing Stereotypes from Representation to Output in Multilingual LLMs
Ariun-Erdene Tumurchuluun, Yusser Al Ghussin, Pinzhen Chen +2
Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects t…
When Tokenization is Secretly Output Supervision
Tanja Baeumel, Josef van Genabith, Simon Ostermann
Tokenization in language models is treated by default as an input preprocessing decision. We argue that this framing is incomplete: in autoregressive models, tokenizer granularity…
Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov +2
Multilingual large language models (mLLMs) achieve strong performance in machine translation, yet our understanding of the mechanisms by which they transform representations from o…