4 papers
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
Zhida He, Xiaoyu Wen, Han Qi +7
Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based multi-turn jailbreak method…
SAJA: A State-Action Joint Attack Framework on Multi-Agent Deep Reinforcement Learning
Weiqi Guo, Guanjun Liu, Ziyuan Zhou
Multi-Agent Deep Reinforcement Learning (MADRL) has shown potential for cooperative and competitive tasks such as autonomous driving and strategic gaming. However, models trained b…
PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning
Weiran Guo, Guanjun Liu, Ziyuan Zhou +1
Reinforcement Learning (RL) is widely used in tasks where agents interact with an environment to maximize rewards. Building on this foundation, Safe Reinforcement Learning (Safe RL…
Partially Observable Mean Field Multi-Agent Reinforcement Learning Based on Graph-Attention
Min Yang, Guanjun Liu, Ziyuan Zhou
Traditional multi-agent reinforcement learning algorithms are difficultly applied in a large-scale multi-agent environment. The introduction of mean field theory has enhanced the s…