2 citations · 2 across the 2 of their papers we have counts for
Showing cs.CRShow all
3 papers · 1 filter
cs.CR2025
NonTextual Target Attack
Xinzhe Huang, Wenjing Hu, Tianhang Zheng +6
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However,…
cs.CR2025
Dynamic Jailbreaking Attack
Kedong Xiu, Yunhan Yang, Churui Zeng +6
Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a static optimization strategy. However, thi…
cs.CR2024★ 2 cited
Releasing Malevolence from Benevolence: The Menace of Benign Data on Machine Unlearning
Binhao Ma, Tianhang Zheng, Hongsheng Hu +5
Machine learning models trained on vast amounts of real or synthetic data often achieve outstanding predictive performance across various domains. However, this utility comes with…