2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.CL2024★ 2 cited
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
Yifei Wang, Dizhan Xue, Shengjie Zhang +1
With the prosperity of large language models (LLMs), powerful LLM-based intelligent agents have been developed to provide customized services with a set of user-defined tools. Stat…
cs.CV2023
Erasing Self-Supervised Learning Backdoor by Cluster Activation Masking
Shengsheng Qian, Dizhan Xue, Yifei Wang +3
Self-Supervised Learning (SSL) is an effective paradigm for learning representations from unlabeled data, such as text, images, and videos. However, researchers have recently found…