Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
Maxime Riché, Daniel Tan, Vili Kohonen +1
Inoculation prompting is a selective-generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), a family of methods that similarly reduce…
cs.AI2026
Implementing surrogate goals for safer bargaining in LLM-based agents
Caspar Oesterheld, Maxime Riché, Filip Sondej +2
Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…