1 paper · 1 filter
Jan Betley, Daniel Tan, Niels Warncke +5
We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting mode…