1 citations · 1 across the 12 of their papers we have counts for
4 papers · 1 filter
Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics
Hangtao Zhang, Yucheng Zhao, Sishun Liu +8
Jailbreak prompts can bypass alignment guardrails in large language models (LLMs) and elicit unsafe outputs, making reliable deployment-time detection critical. Prior detection app…
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
Wei Wan, Yuxuan Ning, Zhicong Huang +7
Federated Learning (FL) is a distributed paradigm aimed at protecting participant data privacy by exchanging model parameters to achieve high-quality model training. However, this…
Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM
Lei Yu, Yechao Zhang, Ziqi Zhou +6
With the rapid development of the Vision-Language Model (VLM), significant progress has been made in Visual Question Answering (VQA) tasks. However, existing VLM often generate ina…
DarkFed: A Data-Free Backdoor Attack in Federated Learning
Minghui Li, Wei Wan, Yuxuan Ning +4
Federated learning (FL) has been demonstrated to be susceptible to backdoor attacks. However, existing academic studies on FL backdoor attacks rely on a high proportion of real cli…