1 paper
Maria Carolina Cornelia Wit, Jun Pang
Recent advances in large language models (LLMs) have raised concerns about jailbreaking attacks, i.e., prompts that bypass safety mechanisms. This paper investigates the use of mul…