1 paper
T. Ben Thompson, Michael Sklar
Many publicly available language models have been safety tuned to reduce the likelihood of toxic or liability-inducing text. To redteam or jailbreak these models for compliance wit…