1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Jiawei Zhang, Andrew Estornell, David D. Baek +2
Large Language Models (LLMs) exhibit strong but shallow alignment: they directly refuse harmful queries when a refusal is expected at the very start of an assistant turn, yet this…