activity
20242026
collaborators

5 papers

cs.AI2026

Implementing surrogate goals for safer bargaining in LLM-based agents

Caspar Oesterheld, Maxime Riché, Filip Sondej +2

Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…

cs.LG2025

Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning

Filip Sondej, Yushi Yang

Current unlearning and safety training methods consistently fail to remove dangerous knowledge from language models. We identify the root cause - unlearning targets representations…

cs.LG2025

Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization

Filip Sondej, Yushi Yang, Mikołaj Kniejski +1

Language models can retain dangerous knowledge and skills even after extensive safety fine-tuning, posing both misuse and misalignment risks. Recent studies show that even speciali…

cs.AI2025

Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems

Pierre Peigne-Lefebvre, Mikolaj Kniejski, Filip Sondej +4

As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agent…

cs.LG2024

How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

Yushi Yang, Filip Sondej, Harry Mayne +2

Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of alg…