1 citations · 2 across the 6 of their papers we have counts for
4 papers · 1 filter
NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
Chuhan Zhang, Ye Zhang, Bowen Shi +5
Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechan…
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
Jiacheng Liang, Zian Wang, Lauren Hong +2
Various watermarking methods (``watermarkers'') have been proposed to identify LLM-generated texts; yet, due to the lack of unified evaluation platforms, many critical questions re…
Watch the Watcher! Backdoor Attacks on Security-Enhancing Diffusion Models
Changjiang Li, Ren Pang, Bochuan Cao +4
Thanks to their remarkable denoising capabilities, diffusion models are increasingly being employed as defensive tools to reinforce the security of other models, notably in purifyi…
On the Difficulty of Defending Contrastive Learning against Backdoor Attacks
Changjiang Li, Ren Pang, Bochuan Cao +4
Recent studies have shown that contrastive learning, like supervised learning, is highly vulnerable to backdoor attacks wherein malicious functions are injected into target models,…