From the 1 of 8 linked papers with an AI index.
8 papers
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
Maxime Riché, Daniel Tan, Vili Kohonen +1
The paper proposes inoculation adapters, a LoRA‑based method that trains on undesired traits and then discards the adapter to improve selective generalization of desired capabiliti…
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
Jan DubiÅski, Jan Betley, Anna Sztyber-Betley +2
Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregio…
Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
Chloe Li, Mary Phuong, Daniel Tan
As AI systems become more capable of complex agentic tasks, they also become more capable of pursuing undesirable objectives and causing harm. Previous work has attempted to catch…
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Daniel Tan, Anders Woodruff, Niels Warncke +4
Language model finetuning often results in learning undesirable traits in combination with desired ones. To address this, we propose inoculation prompting: modifying finetuning dat…
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
Jan Wehner, Sahar Abdelnabi, Daniel Tan +2
Representation Engineering (RepE) is a novel paradigm for controlling the behavior of LLMs. Unlike traditional approaches that modify inputs or fine-tune the model, RepE directly m…
LISA Technical Report: An Agentic Framework for Smart Contract Auditing
Izaiah Sun, Daniel Tan, Andy Deng
We present LISA, an agentic smart contract vulnerability detection framework that combines rule-based and logic-based methods to address a broad spectrum of vulnerabilities in smar…