4 papers
PACT: Preserving Anchored Cores in Task-vectors for Model Merging
Ningyuan Shi, Zhipeng Zhou, Hao Wang +2
Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model. Most exi…
Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
Zheng Zhang, Jiarui He, Yuchen Cai +4
As large language model (LLM) agents increasingly automate complex web tasks, they boost productivity while simultaneously introducing new security risks. However, relevant studies…
The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games
Zhang Zheng, Deheng Ye, Peilin Zhao +1
Large language model (LLM) agents have shown remarkable progress in social deduction games (SDGs). However, existing approaches primarily focus on information processing and strate…
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
Zheng Zhang, Peilin Zhao, Deheng Ye +1
Jailbreak attacks aim to exploit large language models (LLMs) by inducing them to generate harmful content, thereby revealing their vulnerabilities. Understanding and addressing th…