12 papers
Data Attribution of Emergent Misalignment with Persona Features
Clemens Vetter, David Kaczér, Lucie Flek +1
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attri…
A Unified Moral-Value Dataset for Instruction Tuning
Zhaohui Zeng, Florian Mai
Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human values is still an open prob…
In-Training Defenses against Emergent Misalignment in Language Models
David Kaczér, Magnus Jørgenvåg, Clemens Vetter +4
Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EM): Even a small, domain-specific fin…
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
Amr Moustafa, Max Feser, Florian Mai
Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain…
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
Magnus Jørgenvåg, David Kaczér, Lasse Ruttert +3
Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While EM has been extensive…
Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
Shivam Rawat, Lucie Flek, Florian Mai +1
Reasoning in large language models is often discussed as a single capability, but some of its gains may stem from simpler underlying operations. We examine two such primitives, rec…