collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Data Attribution of Emergent Misalignment with Persona Features

Clemens Vetter, David Kaczér, Lucie Flek +1

Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attri…

cs.CL2026

A Unified Moral-Value Dataset for Instruction Tuning

Zhaohui Zeng, Florian Mai

Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human values is still an open prob…

cs.CL2026

Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

Magnus Jørgenvåg, David Kaczér, Lasse Ruttert +3

Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While EM has been extensive…

cs.CL2026

Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?

Shivam Rawat, Lucie Flek, Florian Mai +1

Reasoning in large language models is often discussed as a single capability, but some of its gains may stem from simpler underlying operations. We examine two such primitives, rec…

cs.CL2026

Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi

Shiza Fatimah, Aniket Sen, Sophia Falk +3

The dominance of large multilingual foundation models has widened linguistic inequalities in Natural Language Processing (NLP), often leaving low-resource languages underrepresente…

cs.CL2026

Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models

Christian Nickel, Laura Schrewe, Florian Mai +1

Theory of Mind (ToM) refers to an agent's ability to model the internal states of others. Contributing to the debate whether large language models (LLMs) exhibit genuine ToM capabi…