collaborators

6 papers

cs.CL2026

RepSelect: Robust LLM Unlearning via Representation Selectivity

Filip Sondej, Yushi Yang, Adam Mahdi

When LLM weights are open or fine-tuning is available through an API, suppressing hazardous knowledge and tendencies is not enough: removal has to be deep enough that an adversary…

cs.AI2026

Implementing surrogate goals for safer bargaining in LLM-based agents

Caspar Oesterheld, Maxime Riché, Filip Sondej +2

Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…

cs.LG2025

Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization

Filip Sondej, Yushi Yang, Mikołaj Kniejski +1

Language models can retain dangerous knowledge and skills even after extensive safety fine-tuning, posing both misuse and misalignment risks. Recent studies show that even speciali…

cs.LG2025

Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning

Filip Sondej, Yushi Yang

Current unlearning and safety training methods consistently fail to remove dangerous knowledge from language models. We identify the root cause - unlearning targets representations…

cs.LG2025

How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis

Yushi Yang, Filip Sondej, Harry Mayne +2

Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of alg…

cs.AI2025

Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems

Pierre Peigne-Lefebvre, Mikolaj Kniejski, Filip Sondej +4

As AI agents are increasingly adopted to collaborate on complex objectives, ensuring the security of autonomous multi-agent systems becomes crucial. We develop simulations of agent…