activity
20242026
collaborators

9 papers

cs.CL2026

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

Yinjie Cheng, Paul Youssef, Christin Seifert +2

Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remains the default operation for adapting L…

cs.LG2026

One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

Ali Holmov, Paul Youssef, Nandi Schoots +1

Knowledge editing methods such as ROME and MEMIT update factual associations in transformer models by modifying MLP weights. While evaluated mainly by output behavior, their intern…

cs.CL2026

Tracing and Reversing Edits in LLMs

Paul Youssef, Zhixue Zhao, Christin Seifert +1

Knowledge editing methods (KEs) are a cost-effective way to update the factual content of large language models (LLMs), but they pose a dual-use risk. While KEs are beneficial for…

cs.CL2026

Persuasion Tokens for Editing Factual Knowledge in LLMs

Paul Youssef, Christin Seifert, Jörg Schlötterer

In-context knowledge editing (IKE) is a promising technique for updating Large Language Models (LLMs) with new information. However, IKE relies on lengthy, fact-specific demonstrat…

cs.CL2025

Position: Editing Large Language Models Poses Serious Safety Risks

Paul Youssef, Zhixue Zhao, Daniel Braun +2

Large Language Models (LLMs) contain large amounts of facts about the world. These facts can become outdated over time, which has led to the development of knowledge editing method…

cs.CL2025

How to Make LLMs Forget: On Reversing In-Context Knowledge Edits

Paul Youssef, Zhixue Zhao, Jörg Schlötterer +1

In-context knowledge editing (IKE) enables efficient modification of large language model (LLM) outputs without parameter changes and at zero-cost. However, it can be misused to ma…