1 paper
Zhao Tong, Chunlin Gong, Yiping Zhang +5
From generating headlines to fabricating news, the Large Language Models (LLMs) are typically assessed by their final outputs, under the safety assumption that a refusal response s…