3 papers
cs.CR2025
Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
Gejian Zhao, Hanzhou Wu, Xinpeng Zhang
Backdoor attacks pose a significant security threat to natural language processing (NLP) systems, but existing methods lack explainable trigger mechanisms and fail to quantitativel…
cs.CR2025
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
Gejian Zhao, Hanzhou Wu, Xinpeng Zhang +1
Chain-of-Thought (CoT) enhances an LLM's ability to perform complex reasoning tasks, but it also introduces new security issues. In this work, we present ShadowCoT, a novel backdoo…
cs.CR2024
Transferable Watermarking to Self-supervised Pre-trained Graph Encoders by Trigger Embeddings
Xiangyu Zhao, Hanzhou Wu, Xinpeng Zhang
Recent years have witnessed the prosperous development of Graph Self-supervised Learning (GSSL), which enables to pre-train transferable foundation graph encoders. However, the eas…