1 paper · 1 filter
Zhao Tong, Chunlin Gong, Yiping Zhang +5
From generating headlines to fabricating news, the Large Language Models (LLMs) are typically assessed by their final outputs, under the safety assumption that a refusal response s…