1 paper
Dongdong Zhang, Tengchao Lv, Yilin Jia +9
Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Re…