adapter training 1backdoor mitigation 1bias 1ethical AI 1inoculation prompting 1large language models 1LLM evaluation 1model alignment 1selective generalization 1value leakage 1
From the 2 of 5 linked papers with an AI index.
22 citations · 22 across the 4 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026★ 22 cited
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Jan Betley, Daniel Tan, Niels Warncke +5
We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting mode…
cs.CL2025
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Daniel Tan, Anders Woodruff, Niels Warncke +4
Language model finetuning often results in learning undesirable traits in combination with desired ones. To address this, we propose inoculation prompting: modifying finetuning dat…