adapter training 1backdoor mitigation 1inoculation prompting 1large language models 1selective generalization 1
From the 1 of 8 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
Maxime Riché, Daniel Tan, Vili Kohonen +1
The paper proposes inoculation adapters, a LoRA‑based method that trains on undesired traits and then discards the adapter to improve selective generalization of desired capabiliti…
cs.AI2026
Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
Chloe Li, Mary Phuong, Daniel Tan
As AI systems become more capable of complex agentic tasks, they also become more capable of pursuing undesirable objectives and causing harm. Previous work has attempted to catch…