17 citations · 17 across the 1 of their papers we have counts for
1 paper · 1 filter
Yuan Yuan, Tina Sriskandarajah, Anna-Luisa Brakman +4
Large Language Models used in ChatGPT have traditionally been trained to learn a refusal boundary: depending on the user's intent, the model is taught to either fully comply or out…