1 paper
Jeremy Kritz, Vaughn Robinson, Robert Vacareanu +7
Large Language Models (LLMs) can be used to red team other models (e.g. jailbreaking) to elicit harmful contents. While prior works commonly employ open-weight models or private un…