1 paper · 1 filter
Ran Li, Hao Wang, Chengzhi Mao
Efficient red-teaming method to uncover vulnerabilities in Large Language Models (LLMs) is crucial. While recent attacks often use LLMs as optimizers, the discrete language space m…