3 papers
cs.CR2025
Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
Gejian Zhao, Hanzhou Wu, Xinpeng Zhang
Backdoor attacks pose a significant security threat to natural language processing (NLP) systems, but existing methods lack explainable trigger mechanisms and fail to quantitativel…
cs.CR2025
Yet Another Watermark for Large Language Models
Siyuan Bao, Ying Shi, Zhiguang Yang +2
Existing watermarking methods for large language models (LLMs) mainly embed watermark by adjusting the token sampling prediction or post-processing, lacking intrinsic coupling with…
cs.MM2025
RFNNS: Robust Fixed Neural Network Steganography with Universal Text-to-Image Models
Yu Cheng, Jiuan Zhou, Jiawei Chen +2
With the rapid development of generative AI, image steganography has garnered widespread attention due to its unique concealment. Recent studies have demonstrated the practical adv…