adapter training 1backdoor mitigation 1inoculation prompting 1large language models 1selective generalization 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
Maxime Riché, Daniel Tan, Vili Kohonen +1
The paper proposes inoculation adapters, a LoRA‑based method that trains on undesired traits and then discards the adapter to improve selective generalization of desired capabiliti…
cs.AI2026
Implementing surrogate goals for safer bargaining in LLM-based agents
Caspar Oesterheld, Maxime Riché, Filip Sondej +2
Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…
cs.CL2025
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Daniel Tan, Anders Woodruff, Niels Warncke +4
Language model finetuning often results in learning undesirable traits in combination with desired ones. To address this, we propose inoculation prompting: modifying finetuning dat…