6 papers · 1 filter
JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills
Xiaoyu Wen, Jiajia Li, Zhida He +11
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematica…
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
Zhida He, Xiaoyu Wen, Han Qi +7
Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based multi-turn jailbreak method…
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
Jiajia Li, Xiaoyu Wen, Zhongtian Ma +3
The growing capabilities of large language models (LLMs) have driven their widespread deployment across diverse domains, even in potentially high-risk scenarios. Despite advances i…
MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
Xiaoyu Wen, Zhida He, Han Qi +7
Ensuring robust safety alignment is crucial for Large Language Models (LLMs), yet existing defenses often lag behind evolving adversarial attacks due to their \textbf{reliance on s…
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law
Shanghai AI Lab, :, Yicheng Bao +115
We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framewo…
ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Ziyu Wan, Yunxiang Li, Xiaoyu Wen +8
Recent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking -- enabling models to monitor, evaluate, a…