5 papers
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
Zhenhao Xu, Wenhan Chang, Yichuan Chen +3
Large Reasoning Models (LRMs) improve performance on complex tasks, but they also make safety control harder at deployment time. In black-box settings, defenders cannot modify mode…
Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor
Wenhan Chang, Tianqing Zhu, Ping Xiong +2
Large Language Models (LLMs) are increasingly deployed in settings where Chain-of-Thought (CoT) is interpreted by users. This creates a new safety risk: attackers may manipulate th…
Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
Wenhan Chang, Tianqing Zhu, Yu Zhao +3
In the era of rapid generative AI development, interactions with large language models (LLMs) pose increasing risks of misuse. Prior research has primarily focused on attacks using…
Zero-shot Class Unlearning via Layer-wise Relevance Analysis and Neuronal Path Perturbation
Wenhan Chang, Tianqing Zhu, Ping Xiong +3
In the rapid advancement of artificial intelligence, privacy protection has become crucial, giving rise to machine unlearning. Machine unlearning is a technique that removes specif…
Large Language Models Merging for Enhancing the Link Stealing Attack on Graph Neural Networks
Faqian Guan, Tianqing Zhu, Wenhan Chang +2
Graph Neural Networks (GNNs), specifically designed to process the graph data, have achieved remarkable success in various applications. Link stealing attacks on graph data pose a…