8 papers · 1 filter
Data Attribution of Emergent Misalignment with Persona Features
Clemens Vetter, David Kaczér, Lucie Flek +1
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attri…
A Unified Moral-Value Dataset for Instruction Tuning
Zhaohui Zeng, Florian Mai
Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human values is still an open prob…
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
Magnus Jørgenvåg, David Kaczér, Lasse Ruttert +3
Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While EM has been extensive…
Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
Shivam Rawat, Lucie Flek, Florian Mai +1
Reasoning in large language models is often discussed as a single capability, but some of its gains may stem from simpler underlying operations. We examine two such primitives, rec…
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
Shiza Fatimah, Aniket Sen, Sophia Falk +3
The dominance of large multilingual foundation models has widened linguistic inequalities in Natural Language Processing (NLP), often leaving low-resource languages underrepresente…
Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models
Christian Nickel, Laura Schrewe, Florian Mai +1
Theory of Mind (ToM) refers to an agent's ability to model the internal states of others. Contributing to the debate whether large language models (LLMs) exhibit genuine ToM capabi…