8 papers
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
Ziqing Yang, Rui Wen, Xinlei He +3
Prompt learning is a new machine learning paradigm that has attracted ample attention due to its simplicity and proven efficacy. Despite its growing adoption, the security vulnerab…
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
Rui Wen, Mark Russinovich, Andrew Paverd +2
Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safety- and privacy-critical appli…
Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills
Yiyong Liu, Chia-Yi Hsu, Chun-Ying Huang +3
LLM-powered coding agents increasingly make software supply chain decisions. They generate imports, recommend packages, and write installation commands. Prior work showed that thes…
Peering Behind the Shield: Guardrail Identification in Large Language Models
Ziqing Yang, Yixin Wu, Rui Wen +2
With the rapid adoption of large language models (LLMs), conversational AI agents have become widely deployed across real-world applications. To enhance safety, these agents are of…
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
Ragib Amin Nihal, Rui Wen, Kazuhiro Nakadai +1
Large language models (LLMs) remain vulnerable to multi-turn jailbreaking attacks that exploit conversational context to bypass safety constraints gradually. These attacks target d…
SL-CBM: Enhancing Concept Bottleneck Models with Semantic Locality for Better Interpretability
Hanwei Zhang, Luo Cheng, Rui Wen +3
Explainable AI (XAI) is crucial for building transparent and trustworthy machine learning systems, especially in high-stakes domains. Concept Bottleneck Models (CBMs) have emerged…