1 paper
Alex Irpan, Alexander Matt Turner, Mark Kurzeja +2
An LLM's factuality and refusal training can be compromised by simple changes to a prompt. Models often adopt user beliefs (sycophancy) or satisfy inappropriate requests which are…