activity
20242026
most citedJBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CR2026

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

Zheng Fang, Xiaosen Wang, Shenyi Zhang +2

Jailbreak attacks on audio language models (ALMs) optimize audio perturbations to elicit unsafe generations, and they typically update the entire waveform densely throughout optimi…

cs.CR2025

Selective Masking Adversarial Attack on Automatic Speech Recognition Systems

Zheng Fang, Shenyi Zhang, Tao Wang +3

Extensive research has shown that Automatic Speech Recognition (ASR) systems are vulnerable to audio adversarial attacks. Current attacks mainly focus on single-source scenarios, i…

cs.CR2025★ 1 cited

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation

Shenyi Zhang, Yuchen Zhai, Keyan Guo +7

Despite the implementation of safety alignment strategies, large language models (LLMs) remain vulnerable to jailbreak attacks, which undermine these safety guardrails and pose sig…

cs.CR2024

Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems

Zheng Fang, Tao Wang, Lingchen Zhao +6

In recent years, extensive research has been conducted on the vulnerability of ASR systems, revealing that black-box adversarial example attacks pose significant threats to real-wo…

cs.CR2024

Hijacking Attacks against Neural Networks by Analyzing Training Data

Yunjie Ge, Qian Wang, Huayang Huang +7

Backdoors and adversarial examples are the two primary threats currently faced by deep neural networks (DNNs). Both attacks attempt to hijack the model behaviors with unintended ou…