activity
20242026
collaborators

10 papers

cs.CL2026

Evaluation Awareness in Language Models Has Limited Effect on Behaviour

Amelie Knecht, Lucas Florin, Thilo Hagendorff

Large reasoning models (LRMs) sometimes note in their chain of thought (CoT) that they may be under evaluation. Researchers worry that this verbalised evaluation awareness (VEA) ca…

cs.CL2026

"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior

Roshni Lulla, Fiona Collins, Sanaya Parekh +2

The alignment problem refers to concerns regarding powerful intelligences, ensuring compatibility with human preferences and values as capabilities increase. Current large language…

cs.CL2026

Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment

Laurène Vaugrante, Anietta Weckauff, Thilo Hagendorff

Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misal…

cs.CL2025

Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models

Monika Jotautaitė, Lucius Caviola, David A. Brewster +1

As large language models (LLMs) become more widely deployed, it is crucial to examine their ethical tendencies. Building on research on fairness and discrimination in AI, we invest…

cs.CL2025

Large Reasoning Models Are Autonomous Jailbreak Agents

Thilo Hagendorff, Erik Derner, Nuria Oliver

Jailbreaking -- bypassing built-in safety mechanisms in AI models -- has traditionally required complex technical procedures or specialized human expertise. In this study, we show…

cs.CL2025

On the Inevitability of Left-Leaning Political Bias in Aligned Language Models

Thilo Hagendorff

The guiding principle of AI alignment is to train large language models (LLMs) to be harmless, helpful, and honest (HHH). At the same time, there are mounting concerns that LLMs ex…