13 papers
Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics
Hangtao Zhang, Yucheng Zhao, Sishun Liu +8
Jailbreak prompts can bypass alignment guardrails in large language models (LLMs) and elicit unsafe outputs, making reliable deployment-time detection critical. Prior detection app…
Why Does Little Robustness Help? A Further Step Towards Understanding Adversarial Transferability
Yechao Zhang, Shengshan Hu, Leo Yu Zhang +5
Adversarial examples (AEs) for DNNs have been shown to be transferable: AEs that successfully fool white-box surrogate models can also deceive other black-box models with different…
HiF-DTA: Hierarchical Feature Learning Network for Drug-Target Affinity Prediction
Minghui Li, Yuanhang Wang, Peijin Guo +3
Accurate prediction of Drug-Target Affinity (DTA) is crucial for reducing experimental costs and accelerating early screening in computational drug discovery. While sequence-based…
DarkHash: A Data-Free Backdoor Attack Against Deep Hashing
Ziqi Zhou, Menghao Deng, Yufei Song +6
Benefiting from its superior feature learning capabilities and efficiency, deep hashing has achieved remarkable success in large-scale image retrieval. Recent studies have demonstr…
Towards Real-World Deepfake Detection: A Diverse In-the-wild Dataset of Forgery Faces
Junyu Shi, Minghui Li, Junguo Zuo +8
Deepfakes, leveraging advanced AIGC (Artificial Intelligence-Generated Content) techniques, create hyper-realistic synthetic images and videos of human faces, posing a significant…
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
Wei Wan, Yuxuan Ning, Zhicong Huang +7
Federated Learning (FL) is a distributed paradigm aimed at protecting participant data privacy by exchanging model parameters to achieve high-quality model training. However, this…