4 papers
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
Yu Yan, Sheng Sun, Zenghao Duan +5
Current studies have exposed the risk of Large Language Models (LLMs) generating harmful content by jailbreak attacks. However, they overlook that the direct generation of harmful…
Jailbreak-as-a-Service++: Unveiling Distributed AI-Driven Malicious Information Campaigns Powered by LLM Crowdsourcing
Yu Yan, Sheng Sun, Mingfeng Li +6
To prevent the misuse of Large Language Models (LLMs) for malicious purposes, numerous efforts have been made to develop the safety alignment mechanisms of LLMs. However, as multip…
FNBench: Benchmarking Robust Federated Learning against Noisy Labels
Xuefeng Jiang, Jia Li, Nannan Wu +7
Robustness to label noise within data is a significant challenge in federated learning (FL). From the data-centric perspective, the data quality of distributed datasets can not be…
Na'vi or Knave: Jailbreaking Language Models via Metaphorical Avatars
Yu Yan, Sheng Sun, Junqi Tong +2
Metaphor serves as an implicit approach to convey information, while enabling the generalized comprehension of complex subjects. However, metaphor can potentially be exploited to b…