1 paper
Tanmay Gautam, Alireza Bahramali, Sandeep Atluri
Automated red-teaming methods for large language models typically optimize attack prompts within a fixed, human-designed strategy, leaving the attack strategy itself unchanged. We…